I work across AI agents, cloud-native AI and inference acceleration: Agent Loop / RSI, disaggregated P/D deployments, and tiered KV cache reuse.
The agent layer uses model services; the inference layer connects engines to tiered storage for KV cache reuse; the cloud-native layer provides routing, P/D deployment and compute resources.