<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Yash Agarwal</title><description>MS student in Intelligent Information Systems at Carnegie Mellon&apos;s School of Computer Science, working on machine learning systems.</description><link>https://yash-agarwal.org/</link><item><title>Skipping attention blocks was the easy part</title><link>https://yash-agarwal.org/blog/blasst/</link><guid isPermaLink="true">https://yash-agarwal.org/blog/blasst/</guid><description>BLASST block-sparse attention plus one TMEM address swap: −41.9% attention-kernel time at 95% skip and −20% TTFT at 64K with accuracy above dense — an optimization worklog on a B200.</description><pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate></item><item><title>DSpark on MAX: making speculative decoding beat vLLM</title><link>https://yash-agarwal.org/blog/dspark/</link><guid isPermaLink="true">https://yash-agarwal.org/blog/dspark/</guid><description>How speculative decoding for Gemma4 on Modular&apos;s MAX went from 30% behind vLLM under load to outside its latency/throughput curve on every dataset at concurrency 8 and 64, on one B200.</description><pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate></item><item><title>How I beat NVIDIA at allocating pinned host memory</title><link>https://yash-agarwal.org/blog/pinned-host-memory/</link><guid isPermaLink="true">https://yash-agarwal.org/blog/pinned-host-memory/</guid><description>9× faster than cuMemAllocHost, and faster than CUDA&apos;s own VMM API too — an optimization worklog on an 8×B200 host.</description><pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate></item></channel></rss>