Engineering Leadership: My User Manual
A concise user manual that explains how I collaborate, make technical decisions, and communicate on cross-functional engineering projects.
Leadership Insights
Building teams and driving technical excellence
Published perspectives on AI infrastructure, platform engineering, distributed systems, and technical leadership—focused on conclusions that remain useful beyond one project.
A concise user manual that explains how I collaborate, make technical decisions, and communicate on cross-functional engineering projects.
Building teams and driving technical excellence
A practical equation for turning measured batch throughput, latency limits, replica count, and TPU topology into an ML serving capacity estimate.
Lessons from working across ML infrastructure and GenAI serving: reliable capacity requires a product contract across accelerators, entitlements, admission control, scheduling, and operations.
A concise user manual that explains how I collaborate, make technical decisions, and communicate on cross-functional engineering projects.
A practical way to place LLM safety checks without turning every request into another full model inference.
站在 2025 的尾巴上,聊聊明年我看好的三个方向:Agent 的真正落地,端侧模型的逆袭,以及我们要如何搞定“可信度”这最后一公里。
How “right but useless” comments in meetings waste time, drain focus, and reveal the difference between showmanship and real contribution.
The practical rules I use when automation has to survive retries, manual intervention, and changing infrastructure requirements.
A reflection on the difference between acting for promotions and earning leadership through impact — and how promotion-chasing behaviors harm teams, peers, and the next generation of engineers.
Lessons from building and locally validating a quota-aware LLM gateway: account for work, make reservations explicit, and explain every rejection.
Lessons from leading network capacity automation for ML hardware: model intent explicitly, reserve resources safely, and design for operational reality.
Get notified when I publish new insights on AI infrastructure, platform engineering, and tech leadership.
NO SPAM. UNSUBSCRIBE ANYTIME.
Building scalable systems for large language models
Tools and practices for developer productivity
Patterns for building resilient, scalable systems
Managing teams and technical decisions