Skip to main content

Vincent Li

Technical Leader & Distributed Systems Architect | Reliable GenAI & ML Infrastructure

Last updated: Aug 2026

Summary

Technical leader and distributed systems architect with 10+ years building large-scale infrastructure, including 8+ years across ML and AI systems. Former TL for Vertex AI Provisioned Throughput and GenAI Serving, where I led scaling efforts for fixed-cost LLM capacity and platform reliability. I’m now Tech Lead for Gmail Spam Infrastructure, covering GenAI infrastructure beyond serving—including training, monitoring, and safety—at global scale. I also build open-source GenAI safety and infrastructure across the Aether ecosystem.

Experience

Google

April 2021 - Present | Sunnyvale, CA, US

Tech Lead | GenAI Infra & Capacity Planning & Automation

Workspace GenAI Foundation | Gmail Spam & Abuse Infra
  • Leading the engineering strategy for Gmail's anti-abuse infrastructure, safeguarding billions of users from spam, phishing, and malware.
  • Modernizing the large-scale distributed serving platform for spam classification, optimizing for low-latency and high-throughput processing of global email traffic.
  • Integrating advanced LLM-based detection capabilities into the spam filtering pipeline, improving detection efficacy for emerging threat vectors.
Vertex AI | Provisioned Throughput & GenAI Serving Infra
  • Founded and led Provisioned Throughput — Google's fixed-cost, fixed-term subscription for guaranteed LLM inference serving capacity on Vertex AI. Scaled to $300M+ ARR in year one; established as the backbone infrastructure powering Gemini, Vertex AI, Bard, Search, and Ads.
  • Pioneered quota-tiered LLM inference serving at scale (Vertex AI Native Serving); shipped a CEO-mandated MVP in 30 days, establishing an org-wide model for differentiated service tiers.
  • Implemented end-to-end observability for LLM inference workloads (metrics, logs, traces, events), reducing MTTR by 40% and enabling executive dashboards for platform reliability and cost optimization.
  • Spearheaded cross-org demand forecasting and GPU/TPU allocation for GenAI inference traffic, reducing supply–demand mismatches by 30% and eliminating throttling for top enterprise AI customers.
  • Codified engineering principles & roadmaps (automation, UI–backend contracts, multi-quarter planning), driving platform evolution and boosting team velocity across Gemini infra.
Network Infrastructure | Capacity Delivery
  • Led a cross-functional program delivering on-time, cost-effective, NPI-adaptive network infra automation for ML hardware in Google data centers.
  • Designed and implemented a network switch port reservation system, automating ML machine connectivity; eliminated manual planning errors and secured zero downtime for new ML hardware launches.

State Street

June 2018 - April 2021 | Boston, MA, US

Tech Lead & Architect (AVP) - Machine Learning

Automation & Artificial Intelligence
  • Built intelligent email governance (NER + anomaly detection), achieving 99.99%+ accuracy and eliminating 30+ FTE workload annually, while preventing sensitive data leaks.
  • Developed time-series LSTM models for securities finance, optimizing borrow strategies, and generating $750K savings in Q1.

Panda Electronics

June 2012 - August 2014 | Nanjing, Jiangsu, China

Software Engineer

  • Developed embedded software and backend services for consumer electronics, ensuring reliable delivery at manufacturing scale (millions of units shipped).

Projects

Aether (Safe GenAI Platform)

Architected and built Aether, an end-to-end Safe GenAI Platform integrating four microservices: Atlas (traffic governance), Sentinel (safety supervision), Hyperion (ML inference), and MonitorX (observability). Implements defense-in-depth safety with tiered analysis (heuristics, ML classifiers, LLM-based detection) and quota-managed safety compute budgets.

Atlas

Independently built Atlas, a production-oriented LLM gateway reference implementation with Redis-backed quotas, token-aware rate limiting, priority traffic shaping, dynamic stream reservations, and usage forecasting. Validated locally with 46 automated tests; not operated under sustained production traffic.

Sentinel

Independently built Sentinel, a production-oriented GenAI safety reference implementation with tiered analysis, PII detection/redaction, prompt-injection protection, and toxicity filtering. Previously served through RapidAPI and Cloud Run; validated locally rather than under sustained production traffic.

vLLM (Contributor)

Contributing to core inference serving features and observability enhancements in one of the most widely adopted open-source LLM inference engines.

Hyperion

Independently built Hyperion, a production-oriented ML inference reference implementation covering GPU-aware serving, dynamic batching, Redis caching, Kubernetes autoscaling, Prometheus metrics, tracing, and alerting. A local audit passes 82 tests with 7 remaining failures; no production latency or throughput claims are made.

MonitorX

Independently built MonitorX, a production-oriented ML observability reference implementation with Python SDK instrumentation, drift and resource metrics, multi-channel alerting, FastAPI, InfluxDB, and Streamlit. All 97 non-API tests pass locally; the API test module currently has a stale import that blocks full-suite collection.

Awesome LLM Infra

Authored a curated guidebook on LLM infrastructure, covering the full lifecycle — pre-training, post-training, inference, optimization, and monitoring. Includes references, diagrams, and structured learning paths.

Skills

Infrastructure & Systems

Distributed Systems, High-Availability Architecture, Large-Scale ML/LLM Serving, Capacity Forecasting, Resource Allocation, Cost Optimization

AI & ML Platforms

LLM Inference Serving, AI Platform Architecture, Data Pipelines, NLP, Time-Series Forecasting, Deep Learning

AI Safety & Governance

GenAI Safety Systems, Content Moderation Pipelines, PII Detection & Redaction (Presidio), Prompt Injection Defense, Tiered Safety Analysis Architecture, Safety Compute Budgeting

Reliability & Observability

Telemetry, Incident Management, SLO/SLA Design, CI/CD, Automation Frameworks, Executive Dashboards

Leadership & Collaboration

Technical Vision, Founding TL, Cross-Org Initiatives, Roadmap Planning, Mentorship, Rapid MVP Delivery, Executive Stakeholder Alignment, Open Source Engagement

Education

Rochester Institute of Technology

MS, Computer Science

2015 - 2018

Xidian University

BE, Industrial Design

2008 - 2012

Certifications

  • RPA Lifecycle: Introduction, Discovery and Design

    Automation Anywhere • March 2024

  • Go Essential Training

    LinkedIn • January 2021

  • Deploying Scalable Machine Learning for Data Science

    LinkedIn • June 2020

  • DevOps for Data Scientists

    LinkedIn • June 2020

  • Natural Language Processing

    Microsoft • March 2020

  • Deep Learning Specialization

    DeepLearning.AI • December 2019

  • Enterprise Design Thinking Practitioner

    IBM • August 2019

Awards

  • 5 spot bonuses & 20+ peer bonuses

    Google

  • RIT Coding Competition 3rd place

    Microsoft

  • 2017 NTID Poster Copyright Royalty

    National Technical Institute for the Deaf

  • Chinese Mathematical Olympiad (CMO) 3rd prize

    Chinese Mathematical Society

Volunteering

Freelance Technical Writer

Medium & Zhihu

2019 - Present

Authored technical articles on machine learning and data science, sharing knowledge with a broad developer community.

Organizations

  • Association for Computing Machinery (ACM)

    Professional Member • 2017 - Present

  • Data Science Association (DSA)

    Member • 2019 - Present

Interests

Strategic Gaming

World of Tanks NA Top 10, Brawl Stars NA Rank 1 Trio

Open-source Development

Maintain personal projects, Contribute to community projects