SuperMarioYL

EN  ⇄  中文
Leo — AI systems, made to run in production.

I work across AI agents, cloud-native AI and inference acceleration: Agent Loop / RSI, disaggregated P/D deployments, and tiered KV cache reuse.

System architecture

The agent layer uses model services; the inference layer connects engines to tiered storage for KV cache reuse; the cloud-native layer provides routing, P/D deployment and compute resources.

The agent layer uses model services; the inference layer connects engines to tiered storage for KV cache reuse; the cloud-native layer provides routing, P/D deployment and compute resources.

Capabilities

Agent Loop and RSI, disaggregated Prefill/Decode deployment, and vLLM Connector offload with tiered KV storage.

Tech stack

Tech stack

Parallel tracks

Parallel tracks

Selected work


Happy to talk tech and build together.

blog email github views