A static site connecting a professional profile, bilingual résumé, engineering projects, and long-term topics.
Content direction and design review, with AI-assisted implementation and validation.
Development · Testing · Inquiry
My work spans backend development and test engineering. I focus on business logic, validation, and maintainable systems, and I am extending that experience into AI workflows and LLM inference.
Turning business rules into maintainable services and tools.
Translating requirements into verifiable workflows, with testability and automation in mind.
Exploring AI-assisted testing workflows and studying inference systems.
A static site connecting a professional profile, bilingual résumé, engineering projects, and long-term topics.
Content direction and design review, with AI-assisted implementation and validation.
A traceable learning record for LLM inference, KV cache, and evaluation methods.
Framing study questions, understanding inference mechanisms, and planning reproducible experiments.
Starting with practice. Going deeper.
From testing workflows to inference fundamentals, KV cache, and reproducible experiments.
Python, test automation, domain modeling, and maintainable tools.
理解 LLM 推理中 prefill 和 decode 的区别,以及为什么 prefill 更适合 batching。
梳理 KV Cache 的数据结构、显存估算方式,以及长上下文为什么会放大问题。
分析 Prefix Cache 命中与未命中对首 token 延迟的影响,并记录后续 benchmark 计划。
Find more code and project activity on GitHub.