Fleet-Scale VM Isolation
Directing the transparent migration of untrusted workloads into isolated virtual machines across a multi-million node global fleet with near-zero memory and CPU overhead.
Architecting planet-scale systems at the intersection of virtualization and AI. I build the high-performance, secure infrastructure that powers Google’s most critical workloads, including Gemini, Google DeepMind, and Waymo.
Cumulative infrastructure & engineering cost savings across virtualized & AI/ML fleets
Directing Borg Isolation across Google's multi-million node global compute fleet
NVMe VM hotplugging & heterogeneous CPU/GPU/TPU compute fungibility
+11% ML inference throughput via operator fusion & on-chip 3D graph acceleration
I engineer the foundational infrastructure that powers the modern AI era. As a Tech Lead at Google, I specialize in designing planet-scale systems where uncompromising security meets extreme performance. My work directly enables the reliable, high-speed execution of Google’s most critical workloads, including Gemini, Google DeepMind, and Waymo.
My expertise lies at the bleeding edge of virtualization, systems programming, and high-performance computing. Whether I'm architecting next-generation C++ node runtimes that save over 102+ SWE-years (~$30M+ in cost savings) , directing fleet-wide isolation initiatives across millions of nodes, or crafting patent-pending compiler extensions for 3D deep learning hardware, I thrive on translating highly ambiguous architectural challenges into robust, scalable realities.
I don't just write code; I build resilient engines for scale. Armed with deep expertise in modern C++, distributed systems, and low-level hardware-software co-design, my drive is to push the boundaries of what's possible with today’s hardware to unlock tomorrow's technological leaps.
Directing the transparent migration of untrusted workloads into isolated virtual machines across a multi-million node global fleet with near-zero memory and CPU overhead.
Architecting C++ node runtimes that unify heterogeneous CPU, GPU, and TPU fleets and eliminate multiprocessing training I/O bottlenecks via zero-passthrough NVMe VM hotplugging.
Engineering lock-free multithreaded C++ APIs in Borglet and custom 3D deep learning compiler IR passes featuring operator fusion, memory layout transforms, and INT8 quantization.
High-demand, untrusted, and latency-sensitive workloads executing across millions of fleet nodes.
Transparently converting untrusted jobs into isolated VMs with minimal memory and CPU overhead.
Maximizing global fleet utilization and resolving host-level filesystem I/O bottlenecks.
Systems Infrastructure, Virtualization & AI/ML Runtime
Leading node-level runtime architecture across AI/ML Infrastructure and fleet-wide virtualization frameworks.
Directing a team of 11 engineers to fortify ecosystem security by transparently converting untrusted jobs to execute inside isolated VMs across a multi-million node global fleet, rigorously balancing security with system efficiency.
Architecting and optimizing the future of virtualized execution environments across the fleet to maximize hardware utilization with minimal performance and memory overhead.
Leading node-level runtime architecture for a confidential initiative supporting Google's rapid AI growth, driving cross-functional alignment across Kernel, Networking, and Fleet Management to scale heterogeneous infrastructure.
Spearheaded the architecture and end-to-end implementation of NVMe-based VM hotplugging to plug and unplug NVMe devices without passthrough systems or in-VM dependencies—delivering native host-level filesystem speed and resolving critical I/O bottlenecks for Google DeepMind’s multiprocessing training workloads.
Identified and architected next-generation C++ node runtime and virtualization optimizations enabling seamless ML compute fungibility (CPU/GPU/TPU) to maximize global fleet capacity—driving cumulative infrastructure and engineering cost savings equivalent to 102+ SWE-years (~$30M+ in cost savings) and unlocking scale for Gemini.
Led the architectural design and C++ implementation of a distributed resource management initiative, enabling seamless telemetry extraction and in-place VM updates to execute complex, untrusted multi-tenant ML workloads at planetary scale.
Engineered a high-throughput, multithread-safe C++ API in Borglet for task lifecycle management utilizing lock-free programming to minimize contention. Designed and deployed a novel hierarchical dependency-sharing architecture for modularizing fleet-wide autonomous services within the Borg cluster manager.
Authored a patent-pending C++ compiler extension (pub. #20230237014) for a proprietary IR to support 3D deep learning graphs, unlocking full on-chip hardware acceleration. Boosted ML model inference throughput by 11% via graph-level operator fusion, memory layout transformations, and INT8 mixed-precision optimizations.
Architected a scalable ML experimentation pipeline on GCP and deployed a highly optimized C++/Go ensembling library that boosted a production NLP API’s F1 score by 4.28% and reduced annotation errors by 50%.
Enhanced C++ infrastructure connecting to SpannerDB to capture ML anomaly detection scores, resulting in an 8% increase in bad-actor suspensions and a 10% reduction in scan quota.
Co-invented and authored a C++ compiler extension for a proprietary intermediate representation (IR) supporting 3D deep learning tensor graphs, unlocking full on-chip hardware acceleration and boosting inference throughput by 11%.
Achieved a >56% runtime speedup in a custom C++ Geographic Information System by architecting a concurrent thread-pool engine to parallelize map parsing and pathfinding heuristics, securing a Top-6 placement for route quality.
Designed novel multi-agent Reinforcement Learning algorithms (DQN, PPO) that outperformed state-of-the-art methods by converging to a Nash equilibrium in adversarial, partial-information environments simulating drone swarm behavior.
Developed a Top-5% ranked Reversi AI engine in C. Implemented an aggressively optimized Minimax search featuring alpha-beta pruning, Zobrist transposition tables (memoization), and killer heuristic move ordering under tight latency constraints.
Bachelor of Applied Science in Computer Engineering