Rapid Delivery
30 days
Executive-priority MVP
Shipped quota-tiered LLM serving on a compressed timeline and established a reusable model for differentiated service levels.
View experienceVincent Li · Tech Lead & Distributed Systems Architect
I lead platform architecture across capacity, serving, safety, and reliability—turning ambiguous infrastructure problems into systems teams can operate at scale.
Selected Impact
A snapshot of how I operate across industry delivery, hands-on platform engineering, and technical leadership.
Rapid Delivery
30 days
Shipped quota-tiered LLM serving on a compressed timeline and established a reusable model for differentiated service levels.
View experienceReliability
40%
Built end-to-end observability across metrics, logs, traces, and events to reduce MTTR and make platform health actionable.
View experienceCapacity Planning
30%
Aligned forecasting, accelerator allocation, admission control, and operations around a measurable capacity contract.
Read the approachThree systems that show how I approach safety, traffic governance, and production-scale AI infrastructure.
Solution
An independently built, production-oriented Safe GenAI reference implementation integrating content safety, traffic governance, ML inference, and observability.
Solution
An independently built AI supervision reference implementation for enforcing safety, compliance, and quality policies around LLM applications.
Solution
An independently built, production-oriented LLM traffic and quota gateway using Redis, FastAPI, and Prometheus, validated through a local automated test suite.
Deep dives into the architecture of agentic systems.
Designing the operating system for AI: Durable execution, memory kernels, and cognitive architectures.
Moving from 'Vibe Checks' to metrics. A comprehensive framework for testing stochastic AI systems.
How to design a low-latency, fail-closed security layer for autonomous agents. Lessons from building Guardian.
Writing about engineering, leadership, and AI.
A practical equation for turning measured batch throughput, latency limits, replica count, and TPU topology into an ML serving capacity estimate.
Read ArticleLessons from working across ML infrastructure and GenAI serving: reliable capacity requires a product contract across accelerators, entitlements, admission control, scheduling, and operations.
Read ArticleA concise user manual that explains how I collaborate, make technical decisions, and communicate on cross-functional engineering projects.
Read Article