你好,我是 Johney Zheng,一名算法/开发工程师。
目前专注于 AI Infra 与 CV 算法方向。
这个博客记录我在模型部署、推理优化和工程实践上的笔记。从底层的算子与集群通信, 到上层的推理技术栈,再到相关的算法设计——大多是把踩过的坑和读过的论文整理成自己能复用的形式。
主要方向
AI Infra
推理技术栈、集群通信、Triton Kernel、算子优化,以及长思维链模型对基础设施的影响。
模型部署
离线部署与推理加速,GPU/TPU/XPU 等异构硬件的选型与优化策略。
CV 算法
2D/3D 检测、分割与跟踪,以及相关论文的解读笔记。
工程实践
Python 与 Modern C++、设计模式、并发编程,以及开发环境配置。
提示:按 ⌘K(Windows 为 CtrlK)可以随时全站搜索。
联系
如果想联系我,请优先邮件:zwg0606@gmail.com
或者Follow GitHub @ZhengWG。
Hi, I'm Johney Zheng — an algorithm and software engineer working on AI infrastructure and computer vision.
This blog is where I share my notes and learnings on model deployment, inference optimization, and engineering practice — covering everything from low-level kernel and collective communication, to LLM inference stacks, to algorithmic design. Mostly, it's a place where I turn my mistakes and papers I've read into something I can actually reuse.
What I work on
AI Infrastructure
Inference stacks, collective communication, Triton kernels, operator optimization, and what long chain-of-thought models imply for infra.
Model Deployment
Offline deployment and inference acceleration across heterogeneous targets — GPU, TPU, and other accelerators.
Computer Vision
2D and 3D detection, segmentation, and tracking, plus paper reading notes.
Engineering
Python and modern C++, design patterns, concurrency, and development tooling.
Tip: press ⌘K (or CtrlK) to search the whole site from anywhere.
Contact
Email is the best way to reach me: zwg0606@gmail.com
Or follow GitHub @ZhengWG.