Systems Security & AI Infra · Sunnyvale, CA

Thanos Ariyanayagam

Software Engineer & Tech Lead @ Google

Architecting planet-scale systems at the intersection of virtualization and AI. I build the high-performance, secure infrastructure that powers Google’s most critical workloads, including Gemini, Google DeepMind, and Waymo.

01 / IMPACT
~$30M+ Saved
102+ SWE-Years Equivalent

Cumulative infrastructure & engineering cost savings across virtualized & AI/ML fleets

02 / IMPACT
11 Engineers
Tech Lead Scope (L4)

Directing Borg Isolation across Google's multi-million node global compute fleet

03 / IMPACT
DeepMind & Gemini
Critical AI Workloads

NVMe VM hotplugging & heterogeneous CPU/GPU/TPU compute fungibility

04 / IMPACT
Patent #20230237014
US Published Patent

+11% ML inference throughput via operator fusion & on-chip 3D graph acceleration

01

About Me

I engineer the foundational infrastructure that powers the modern AI era. As a Tech Lead at Google, I specialize in designing planet-scale systems where uncompromising security meets extreme performance. My work directly enables the reliable, high-speed execution of Google’s most critical workloads, including Gemini, Google DeepMind, and Waymo.

My expertise lies at the bleeding edge of virtualization, systems programming, and high-performance computing. Whether I'm architecting next-generation C++ node runtimes that save over 102+ SWE-years (~$30M+ in cost savings) , directing fleet-wide isolation initiatives across millions of nodes, or crafting patent-pending compiler extensions for 3D deep learning hardware, I thrive on translating highly ambiguous architectural challenges into robust, scalable realities.

I don't just write code; I build resilient engines for scale. Armed with deep expertise in modern C++, distributed systems, and low-level hardware-software co-design, my drive is to push the boundaries of what's possible with today’s hardware to unlock tomorrow's technological leaps.

Virtualization & Security /01

Fleet-Scale VM Isolation

Directing the transparent migration of untrusted workloads into isolated virtual machines across a multi-million node global fleet with near-zero memory and CPU overhead.

Borg Isolation TI-VM Runtime Multi-Tenant Security
AI / ML Infrastructure /02

Compute Fungibility & Native I/O

Architecting C++ node runtimes that unify heterogeneous CPU, GPU, and TPU fleets and eliminate multiprocessing training I/O bottlenecks via zero-passthrough NVMe VM hotplugging.

CPU / GPU / TPU NVMe Hotplugging Google DeepMind
Hardware-Software Co-Design /03

Compilers & Lock-Free C++

Engineering lock-free multithreaded C++ APIs in Borglet and custom 3D deep learning compiler IR passes featuring operator fusion, memory layout transforms, and INT8 quantization.

Modern C++23 Lock-Free Concurrency Custom Compiler IR
Node Runtime Architecture Stack
Silicon → Hypervisor → Fleet Orchestration
L3 · WORKLOADS

Critical Production AI & Autonomous Systems

Planetary Scale
Gemini Training & Serving
DeepMind Multiprocessing
Waymo Simulation & ML

High-demand, untrusted, and latency-sensitive workloads executing across millions of fleet nodes.

L2 · ISOLATION & RUNTIME

Borg Isolation & TI-VM Next-Gen Execution

Tech Lead · 11 Engineers
Transparent VM Sandboxing
TI-VM Low-Overhead VMM
Lock-Free Borglet C++ API

Transparently converting untrusted jobs into isolated VMs with minimal memory and CPU overhead.

L1 · COMPUTE & STORAGE

Compute Fungibility & Native NVMe Hotplugging

102+ SWE-Yrs (~$30M+ Saved)
CPU / GPU / TPU Fungibility
Native NVMe VM Hotplug
3D DL Compiler IR (Patent)

Maximizing global fleet utilization and resolving host-level filesystem I/O bottlenecks.

02

Experience

2021 — Present

Google

· Sunnyvale, CA

Systems Infrastructure, Virtualization & AI/ML Runtime

Aug 2024 — Present

Software Engineer III & Tech Lead

L4
Nov 2025 — Present

Leading node-level runtime architecture across AI/ML Infrastructure and fleet-wide virtualization frameworks.

Tech Lead — Borg Isolation Commitment 11 Engineers · Multi-Million Node Fleet

Directing a team of 11 engineers to fortify ecosystem security by transparently converting untrusted jobs to execute inside isolated VMs across a multi-million node global fleet, rigorously balancing security with system efficiency.

Tech Lead — TI-VM Next-Gen Runtime Virtualized Execution

Architecting and optimizing the future of virtualized execution environments across the fleet to maximize hardware utilization with minimal performance and memory overhead.

Tech Lead (Node) — AI/ML Infrastructure Kernel · Networking · Fleet Mgmt

Leading node-level runtime architecture for a confidential initiative supporting Google's rapid AI growth, driving cross-functional alignment across Kernel, Networking, and Fleet Management to scale heterogeneous infrastructure.

NVMe-Based VM Hotplugging Google DeepMind · Native I/O

Spearheaded the architecture and end-to-end implementation of NVMe-based VM hotplugging to plug and unplug NVMe devices without passthrough systems or in-VM dependencies—delivering native host-level filesystem speed and resolving critical I/O bottlenecks for Google DeepMind’s multiprocessing training workloads.

Software Engineer II

L3
Aug 2024 — Nov 2025
ML Compute Fungibility & Fleet Virtualization Optimization 102+ SWE-Years · ~$30M+ Cost Savings

Identified and architected next-generation C++ node runtime and virtualization optimizations enabling seamless ML compute fungibility (CPU/GPU/TPU) to maximize global fleet capacity—driving cumulative infrastructure and engineering cost savings equivalent to 102+ SWE-years (~$30M+ in cost savings) and unlocking scale for Gemini.

Distributed VM Resource Management Parity C++ · Planetary Scale

Led the architectural design and C++ implementation of a distributed resource management initiative, enabling seamless telemetry extraction and in-place VM updates to execute complex, untrusted multi-tenant ML workloads at planetary scale.

Earlier Engineering Internships (4x)

Google (3x) · Intel (1x)

Software Engineering Intern

@ Google Borglet · Lock-Free C++
May 2023 — Aug 2023 · Sunnyvale, CA

Engineered a high-throughput, multithread-safe C++ API in Borglet for task lifecycle management utilizing lock-free programming to minimize contention. Designed and deployed a novel hierarchical dependency-sharing architecture for modularizing fleet-wide autonomous services within the Borg cluster manager.

C++ Borglet Lock-Free Concurrency Distributed Systems

Software Engineer Intern

@ Intel Corporation Patent Pub. #20230237014 · +11% Throughput
Sept 2022 — Apr 2023 · Toronto, ON

Authored a patent-pending C++ compiler extension (pub. #20230237014) for a proprietary IR to support 3D deep learning graphs, unlocking full on-chip hardware acceleration. Boosted ML model inference throughput by 11% via graph-level operator fusion, memory layout transformations, and INT8 mixed-precision optimizations.

C++ Compiler Optimization 3D Deep Learning INT8 Quantization

Software Developer Intern

@ Google +4.28% F1 · -50% Errors
May 2022 — Aug 2022 · Waterloo, ON

Architected a scalable ML experimentation pipeline on GCP and deployed a highly optimized C++/Go ensembling library that boosted a production NLP API’s F1 score by 4.28% and reduced annotation errors by 50%.

C++ Go GCP NLP Infrastructure

STEP Intern

@ Google +8% Detection · -10% Scan Quota
May 2021 — Aug 2021 · Waterloo, ON

Enhanced C++ infrastructure connecting to SpannerDB to capture ML anomaly detection scores, resulting in an 8% increase in bad-actor suspensions and a 10% reduction in scan quota.

C++ SpannerDB Anomaly Detection
03

Patents & Selected Work

Proprietary IR & Hardware Acceleration US Patent 2023/0237014 A1

3D Deep Learning Compiler Extension

View Patent Publication

Co-invented and authored a C++ compiler extension for a proprietary intermediate representation (IR) supporting 3D deep learning tensor graphs, unlocking full on-chip hardware acceleration and boosting inference throughput by 11%.

  • Graph-level operator fusion & memory layout transformations
  • Mixed-precision INT8 quantization pass for 3D volumetric convolutions
C++ Compiler IR Operator Fusion INT8 Hardware Co-Design
Multithreaded Spatial Engine & Pathfinding >56% Faster · Top-6 Route Quality

OpenStreetMaps Concurrent GIS

Achieved a >56% runtime speedup in a custom C++ Geographic Information System by architecting a concurrent thread-pool engine to parallelize map parsing and pathfinding heuristics, securing a Top-6 placement for route quality.

  • Asynchronous task scheduling via custom C++ thread pools
  • Thread-safe simulated annealing, 2-opt, and greedy multi-stop heuristics
Modern C++ Multithreading Simulated Annealing A* / Dijkstra
Adversarial Multi-Agent Reinforcement Learning Nash Equilibrium Convergence

NEPIADA Autonomous Drone Swarms

Designed novel multi-agent Reinforcement Learning algorithms (DQN, PPO) that outperformed state-of-the-art methods by converging to a Nash equilibrium in adversarial, partial-information environments simulating drone swarm behavior.

  • Decentralized execution under partial observability constraints
  • Game-theoretic Nash equilibrium convergence in multi-agent arenas
Python PyTorch RLlib Multi-Agent DQN & PPO
Zero-Allocation Game Tree Search in C Top-5% Ranked Engine

Tournament Reversi AI Engine

Developed a Top-5% ranked Reversi AI engine in C. Implemented an aggressively optimized Minimax search featuring alpha-beta pruning, Zobrist transposition tables (memoization), and killer heuristic move ordering under tight latency constraints.

  • Bitboard state representation with O(1) legal move generation
  • Zobrist hashing transposition tables & dynamic endgame solver
C Alpha-Beta Pruning Transposition Tables Bitboards
04

Technical Arsenal

SYS_01

Languages

Modern C++ (17/20/23) C Go Python SQL Bash
INF_02

Infrastructure & Hardware

Borg NVMe SmartNICs Kubernetes gRPC & Protocol Buffers Spanner & BigQuery GCP & AWS Docker Bazel & CMake
PRF_03

Performance & Profiling

Perf GDB Valgrind Google Test (gtest) PyTest CI/CD Pipelines
DOM_04

AI / ML & Core Domains

Systems Design Lock-Free Programming Compiler Optimization Distributed Systems High-Performance Networking PyTorch & TensorFlow Model Pruning / INT8 RLlib
05

Education

2019 — 2024

University of Toronto

· Toronto, ON

Bachelor of Applied Science in Computer Engineering

Sept 2019 — Apr 2024
High Honours (87.5%+ Cumulative Average) Certificate in Artificial Intelligence Minor in Engineering Business
Copied to clipboard