I'm an Applied AI Engineer at Atomicwork, building a runtime auditor for production AI agent runs. Before and alongside that, the performance-engineering side of LLM inference: speculative decoding, KV-cache, quantization, and the systems around them, with three merged open-source PRs to SGLang. MS Data Science, Stony Brook University (2026). There is also a creative side: music, DJ, video editing, 3D and VFX.
Service: artifact evaluation committee, IEEE HPCA 2026.
A Guide to Large Language Model Systems — inference, hardware, retrieval, agents, and safety (June 2026, 243 pages, CC BY 4.0). A guide to how LLM systems actually work, from a single forward pass up to a production system with retrieval, agents, and the safety problems that come with them. It follows one ordinary chat request all the way down the stack, in eight parts: the model, inference, the hardware, serving, knowledge, agency, the system, and safety. Read on Medium · PDF · Source.
Music and DJ (DREAMS, debut single, on Spotify), video editing, 3D and VFX, motion and animation. Full DJ set, DJ LAZER · Mix #1, on YouTube; shorter clips on Instagram.