Skip to main content

Vincent Li · Tech Lead & Distributed Systems Architect

Technical leadership for reliable AI infrastructure.

I lead platform architecture across capacity, serving, safety, and reliability—turning ambiguous infrastructure problems into systems teams can operate at scale.

Selected Impact

Evidence before adjectives.

A snapshot of how I operate across industry delivery, hands-on platform engineering, and technical leadership.

Rapid Delivery

30 days

Executive-priority MVP

Shipped quota-tiered LLM serving on a compressed timeline and established a reusable model for differentiated service levels.

View experience

Reliability

40%

Faster incident recovery

Built end-to-end observability across metrics, logs, traces, and events to reduce MTTR and make platform health actionable.

View experience

Capacity Planning

30%

Fewer supply-demand mismatches

Aligned forecasting, accelerator allocation, admission control, and operations around a measurable capacity contract.

Read the approach

Featured Projects

Three systems that show how I approach safety, traffic governance, and production-scale AI infrastructure.

Aether

Reference ImplementationPlatform

Solution

An independently built, production-oriented Safe GenAI reference implementation integrating content safety, traffic governance, ML inference, and observability.

My Contribution
Defined the shared request lifecycle and integrated Atlas, Sentinel, Hyperion, and MonitorX behind explicit service and deployment contracts.
Evidence
Local integration testing
Read Case Study

Sentinel

Reference ImplementationAI SecurityFormer Public API

Solution

An independently built AI supervision reference implementation for enforcing safety, compliance, and quality policies around LLM applications.

My Contribution
Designed and implemented the tiered inspection pipeline, verdict contract, API packaging, former hosted deployment, and browser simulation.
Evidence
Local functional testing; formerly served via RapidAPI
Read Case Study

Atlas

Reference ImplementationLLM GatewayDistributed Systems

Solution

An independently built, production-oriented LLM traffic and quota gateway using Redis, FastAPI, and Prometheus, validated through a local automated test suite.

My Contribution
Designed and implemented the FastAPI gateway, Redis-backed controls, routing behavior, metrics, and local automated test suite.
Evidence
46 automated tests passed locally

Latest Insights

Writing about engineering, leadership, and AI.

12 min read

From Batch Size to TPU Topology: A Capacity Equation for ML Serving

A practical equation for turning measured batch throughput, latency limits, replica count, and TPU topology into an ML serving capacity estimate.

Read Article
13 min read

GenAI Capacity Is a Product, Not Just an Accelerator Pool

Lessons from working across ML infrastructure and GenAI serving: reliable capacity requires a product contract across accelerators, entitlements, admission control, scheduling, and operations.

Read Article
7 min read

Engineering Leadership: My User Manual

A concise user manual that explains how I collaborate, make technical decisions, and communicate on cross-functional engineering projects.

Read Article