Vincent Li
Technical Leader & Distributed Systems Architect | Reliable GenAI & ML Infrastructure
Summary
Technical leader and distributed systems architect with 10+ years building large-scale infrastructure, including 8+ years across ML and AI systems. Former TL for Vertex AI Provisioned Throughput and GenAI Serving, where I led scaling efforts for fixed-cost LLM capacity and platform reliability. I’m now Tech Lead for Gmail Spam Infrastructure, covering GenAI infrastructure beyond serving—including training, monitoring, and safety—at global scale. I also build open-source GenAI safety and infrastructure across the Aether ecosystem.
Experience
Tech Lead | GenAI Infra & Capacity Planning & Automation
Workspace GenAI Foundation | Gmail Spam & Abuse Infra
- Leading the engineering strategy for Gmail's anti-abuse infrastructure, safeguarding billions of users from spam, phishing, and malware.
- Modernizing the large-scale distributed serving platform for spam classification, optimizing for low-latency and high-throughput processing of global email traffic.
- Integrating advanced LLM-based detection capabilities into the spam filtering pipeline, improving detection efficacy for emerging threat vectors.
Vertex AI | Provisioned Throughput & GenAI Serving Infra
- Founded and led Provisioned Throughput — Google's fixed-cost, fixed-term subscription for guaranteed LLM inference serving capacity on Vertex AI. Scaled to $300M+ ARR in year one; established as the backbone infrastructure powering Gemini, Vertex AI, Bard, Search, and Ads.
- Pioneered quota-tiered LLM inference serving at scale (Vertex AI Native Serving); shipped a CEO-mandated MVP in 30 days, establishing an org-wide model for differentiated service tiers.
- Implemented end-to-end observability for LLM inference workloads (metrics, logs, traces, events), reducing MTTR by 40% and enabling executive dashboards for platform reliability and cost optimization.
- Spearheaded cross-org demand forecasting and GPU/TPU allocation for GenAI inference traffic, reducing supply–demand mismatches by 30% and eliminating throttling for top enterprise AI customers.
- Codified engineering principles & roadmaps (automation, UI–backend contracts, multi-quarter planning), driving platform evolution and boosting team velocity across Gemini infra.
Network Infrastructure | Capacity Delivery
- Led a cross-functional program delivering on-time, cost-effective, NPI-adaptive network infra automation for ML hardware in Google data centers.
- Designed and implemented a network switch port reservation system, automating ML machine connectivity; eliminated manual planning errors and secured zero downtime for new ML hardware launches.
State Street
June 2018 - April 2021 | Boston, MA, USTech Lead & Architect (AVP) - Machine Learning
Automation & Artificial Intelligence
- Built intelligent email governance (NER + anomaly detection), achieving 99.99%+ accuracy and eliminating 30+ FTE workload annually, while preventing sensitive data leaks.
- Developed time-series LSTM models for securities finance, optimizing borrow strategies, and generating $750K savings in Q1.
Panda Electronics
June 2012 - August 2014 | Nanjing, Jiangsu, ChinaSoftware Engineer
- Developed embedded software and backend services for consumer electronics, ensuring reliable delivery at manufacturing scale (millions of units shipped).
Projects
Architected and built Aether, an end-to-end Safe GenAI Platform integrating four microservices: Atlas (traffic governance), Sentinel (safety supervision), Hyperion (ML inference), and MonitorX (observability). Implements defense-in-depth safety with tiered analysis (heuristics, ML classifiers, LLM-based detection) and quota-managed safety compute budgets.
Independently built Atlas, a production-oriented LLM gateway reference implementation with Redis-backed quotas, token-aware rate limiting, priority traffic shaping, dynamic stream reservations, and usage forecasting. Validated locally with 46 automated tests; not operated under sustained production traffic.
Independently built Sentinel, a production-oriented GenAI safety reference implementation with tiered analysis, PII detection/redaction, prompt-injection protection, and toxicity filtering. Previously served through RapidAPI and Cloud Run; validated locally rather than under sustained production traffic.
Contributing to core inference serving features and observability enhancements in one of the most widely adopted open-source LLM inference engines.
Independently built Hyperion, a production-oriented ML inference reference implementation covering GPU-aware serving, dynamic batching, Redis caching, Kubernetes autoscaling, Prometheus metrics, tracing, and alerting. A local audit passes 82 tests with 7 remaining failures; no production latency or throughput claims are made.
Independently built MonitorX, a production-oriented ML observability reference implementation with Python SDK instrumentation, drift and resource metrics, multi-channel alerting, FastAPI, InfluxDB, and Streamlit. All 97 non-API tests pass locally; the API test module currently has a stale import that blocks full-suite collection.
Authored a curated guidebook on LLM infrastructure, covering the full lifecycle — pre-training, post-training, inference, optimization, and monitoring. Includes references, diagrams, and structured learning paths.
Skills
Distributed Systems, High-Availability Architecture, Large-Scale ML/LLM Serving, Capacity Forecasting, Resource Allocation, Cost Optimization
LLM Inference Serving, AI Platform Architecture, Data Pipelines, NLP, Time-Series Forecasting, Deep Learning
GenAI Safety Systems, Content Moderation Pipelines, PII Detection & Redaction (Presidio), Prompt Injection Defense, Tiered Safety Analysis Architecture, Safety Compute Budgeting
Telemetry, Incident Management, SLO/SLA Design, CI/CD, Automation Frameworks, Executive Dashboards
Technical Vision, Founding TL, Cross-Org Initiatives, Roadmap Planning, Mentorship, Rapid MVP Delivery, Executive Stakeholder Alignment, Open Source Engagement
Education
Rochester Institute of Technology
MS, Computer Science
2015 - 2018
Xidian University
BE, Industrial Design
2008 - 2012
Certifications
-
RPA Lifecycle: Introduction, Discovery and Design
Automation Anywhere • March 2024
-
Go Essential Training
LinkedIn • January 2021
-
Deploying Scalable Machine Learning for Data Science
LinkedIn • June 2020
-
DevOps for Data Scientists
LinkedIn • June 2020
-
Natural Language Processing
Microsoft • March 2020
-
Deep Learning Specialization
DeepLearning.AI • December 2019
-
Enterprise Design Thinking Practitioner
IBM • August 2019
Awards
-
5 spot bonuses & 20+ peer bonuses
Google
-
RIT Coding Competition 3rd place
Microsoft
-
2017 NTID Poster Copyright Royalty
National Technical Institute for the Deaf
-
Chinese Mathematical Olympiad (CMO) 3rd prize
Chinese Mathematical Society
Volunteering
Freelance Technical Writer
Medium & Zhihu
2019 - Present
Authored technical articles on machine learning and data science, sharing knowledge with a broad developer community.
Organizations
-
Association for Computing Machinery (ACM)
Professional Member • 2017 - Present
-
Data Science Association (DSA)
Member • 2019 - Present
Interests
Strategic Gaming
World of Tanks NA Top 10, Brawl Stars NA Rank 1 Trio
Open-source Development
Maintain personal projects, Contribute to community projects