I'm an AI researcher focused on large language model inference, transformer memory systems, and efficient long-context evaluation. I build research systems that bridge algorithmic ideas with practical GPU and inference engineering, with an emphasis on reproducibility and open implementations.
My current research includes DKV (Differential KV Cache Compression), a training-free approach to reducing KV-cache memory through anchor representations, low-rank differential reconstruction, and exact residual preservation. DKV combines compression techniques with optimized inference runtimes and is currently under review at IEEE Transactions on Emerging Topics in Computational Intelligence (TETCI).
I'm also developing CRBench, a method-agnostic and resource-aware framework for evaluating long-context LLM systems. It evaluates contextual capability retention against dense reference models while accounting for memory efficiency and inference performance. The current release is an initial research release, with a full version being prepared for TMLR submission.
Beyond these projects, I work on LLM inference optimization, GPU systems, KV-cache architectures, long-context evaluation, model serving, and AI benchmarking. I enjoy taking ideas from first principles, validating them experimentally, and turning them into systems that others can reproduce and extend.
Currently exploring adaptive KV-cache compression, memory-efficient transformer inference, long-context evaluation, and GPU-optimized LLM serving.
I believe research is more useful when it can be reproduced, inspected, and extended. My projects aim to provide complete implementations, benchmarks, documentation, and reproducible experiments wherever possible.
- ResearchGate: https://www.researchgate.net/profile/Om-Chimurkar
- Google Scholar: https://scholar.google.com/citations?user=7NXAM-wAAAAJ
- ORCID: https://orcid.org/0009-0004-0518-4598
- LinkedIn: https://www.linkedin.com/in/om-chimurkar
- Email: omchimurkar45@gmail.com




