<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Quantization on Harvard Systems Group</title><link>https://systems.seas.harvard.edu/tags/quantization/</link><description>Recent content in Quantization on Harvard Systems Group</description><generator>Hugo</generator><language>en</language><copyright>Copyright © 2026 The President and Fellows of Harvard College</copyright><lastBuildDate>Mon, 27 Jul 2026 11:30:00 -0400</lastBuildDate><atom:link href="https://systems.seas.harvard.edu/tags/quantization/index.xml" rel="self" type="application/rss+xml"/><item><title>MorphServe: Making Model Precision Elastic for Bursty LLM Serving</title><link>https://systems.seas.harvard.edu/blog/make-model-precision-elastic/</link><pubDate>Mon, 27 Jul 2026 11:30:00 -0400</pubDate><guid>https://systems.seas.harvard.edu/blog/make-model-precision-elastic/</guid><description>&lt;p&gt;Most LLM servers treat a deployed model as fixed: its weights have one
precision, its KV cache has one memory budget, and both remain unchanged until
the process restarts.&lt;/p&gt;</description></item></channel></rss>