<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>KVM on Charlie Chiang's blog</title><link>https://charlie0129.github.io/blog/tags/kvm/</link><description>Recent content in KVM on Charlie Chiang's blog</description><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Tue, 06 Oct 2026 22:17:00 +0800</lastBuildDate><atom:link href="https://charlie0129.github.io/blog/tags/kvm/index.xml" rel="self" type="application/rss+xml"/><item><title>Making VMs Give Memory Back to the Host Like Containers Do</title><link>https://charlie0129.github.io/blog/p/vm-memory-give-back/</link><pubDate>Tue, 06 Oct 2026 22:17:00 +0800</pubDate><guid>https://charlie0129.github.io/blog/p/vm-memory-give-back/</guid><description>&lt;h2 id="the-problem">The Problem
&lt;/h2>&lt;p>Containers are cheap on memory because they share the host kernel. When a process exits, its pages go straight back to the host&amp;rsquo;s free list. A VM is not like that. On my Proxmox VE hosts, a guest that touched 4 GiB once keeps 4 GiB of host memory forever, even if the guest is now idle with almost nothing running. The guest kernel knows the pages are free. The host does not.&lt;/p>
&lt;p>Then I noticed OrbStack, the Docker Desktop replacement on macOS, does not behave this way. OrbStack runs all containers in a single Linux VM, yet when you stop a container, the macOS memory of that VM drops within seconds. So &amp;ldquo;VMs do not give memory back&amp;rdquo; is a choice, not a law. I wanted to know how OrbStack does it and how much of it I can get on Proxmox VE.&lt;/p>
&lt;p>Everything below was measured on a Proxmox VE 9.1 host with an Alpine test guest, and on an OrbStack 2.2.3 VM, with a root shell obtained the usual way:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-bash" data-lang="bash">&lt;span class="line">&lt;span class="cl">docker run -it --rm --privileged --pid&lt;span class="o">=&lt;/span>host justincormack/nsenter1
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;h2 id="why-guests-hold-on-to-memory">Why Guests Hold On to Memory
&lt;/h2>&lt;p>Guest RAM is anonymous memory in the QEMU process. The first time the guest touches a page, KVM faults it in and the host backs it. When the guest later frees that page, it just moves to the guest&amp;rsquo;s free list. Nothing tells the host, so the host page stays resident. Three mechanisms can change this:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Classic ballooning&lt;/strong> is host-driven. Proxmox&amp;rsquo;s &lt;code>pvestatd&lt;/code> only inflates the balloon when the host is above about 80% memory, and only down to the VM&amp;rsquo;s minimum. Nothing happens when a guest process exits.&lt;/li>
&lt;li>&lt;strong>KSM&lt;/strong> deduplicates identical pages. Useful, but it does not return freed memory.&lt;/li>
&lt;li>&lt;strong>Free page reporting&lt;/strong> is guest-driven. When the guest frees a large contiguous block, the virtio-balloon driver reports it to the host and QEMU calls &lt;code>madvise(MADV_DONTNEED)&lt;/code> on the range. The host RSS drops within a couple of seconds. This has existed since QEMU 5.1 and Linux 5.7 as the &lt;code>free-page-reporting=on&lt;/code> property of the balloon device.&lt;/li>
&lt;/ul>
&lt;p>So the mechanism is there. The question is why it was not working for me.&lt;/p>
&lt;h2 id="proxmox-already-enables-it-unless-you-turned-it-off">Proxmox Already Enables It, Unless You Turned It Off
&lt;/h2>&lt;p>It turns out Proxmox VE&amp;rsquo;s &lt;code>qemu-server&lt;/code> adds &lt;code>free-page-reporting=on&lt;/code> to the balloon device automatically, for any VM with ballooning enabled and machine version 6.2 or newer:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;span class="lnt">2
&lt;/span>&lt;span class="lnt">3
&lt;/span>&lt;span class="lnt">4
&lt;/span>&lt;span class="lnt">5
&lt;/span>&lt;span class="lnt">6
&lt;/span>&lt;span class="lnt">7
&lt;/span>&lt;span class="lnt">8
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-perl" data-lang="perl">&lt;span class="line">&lt;span class="cl">&lt;span class="c1"># /usr/share/perl5/PVE/QemuServer.pm&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="c1"># enable balloon by default, unless explicitly disabled&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">if&lt;/span> &lt;span class="p">(&lt;/span>&lt;span class="o">!&lt;/span>&lt;span class="nb">defined&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="nv">$conf&lt;/span>&lt;span class="o">-&amp;gt;&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="n">balloon&lt;/span>&lt;span class="p">})&lt;/span> &lt;span class="o">||&lt;/span> &lt;span class="nv">$conf&lt;/span>&lt;span class="o">-&amp;gt;&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="n">balloon&lt;/span>&lt;span class="p">})&lt;/span> &lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">my&lt;/span> &lt;span class="nv">$pciaddr&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="n">print_pci_addr&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="s">&amp;#34;balloon0&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="nv">$bridges&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="nv">$arch&lt;/span>&lt;span class="p">);&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">my&lt;/span> &lt;span class="nv">$ballooncmd&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="s">&amp;#34;virtio-balloon-pci,id=balloon0$pciaddr&amp;#34;&lt;/span>&lt;span class="p">;&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nv">$ballooncmd&lt;/span> &lt;span class="o">.=&lt;/span> &lt;span class="s">&amp;#34;,free-page-reporting=on&amp;#34;&lt;/span> &lt;span class="k">if&lt;/span> &lt;span class="n">min_version&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="nv">$machine_version&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="mi">6&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="mi">2&lt;/span>&lt;span class="p">);&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nb">push&lt;/span> &lt;span class="nv">@$devices&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="s">&amp;#39;-device&amp;#39;&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="nv">$ballooncmd&lt;/span>&lt;span class="p">;&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>My test guest had &lt;code>balloon: 0&lt;/code> in its config. That does not just disable automatic ballooning. It removes the balloon device entirely, and free page reporting goes with it. I had set it on several VMs over the years because &amp;ldquo;I don&amp;rsquo;t want the host to steal memory from this VM&amp;rdquo;. Oops.&lt;/p>
&lt;p>Here is what that one line costs. The guest allocates and touches 1.2 GiB, holds it for 20 seconds, frees it, and I watch the kvm process RSS on the host:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Config&lt;/th>
&lt;th>Guest frees 1.2 GiB&lt;/th>
&lt;th>Host RSS&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>balloon: 0&lt;/code>&lt;/td>
&lt;td>yes&lt;/td>
&lt;td>237 → 1447 → &lt;strong>1447 MiB&lt;/strong>, never returns&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>balloon default&lt;/td>
&lt;td>yes&lt;/td>
&lt;td>243 → 1451 → &lt;strong>251 MiB&lt;/strong>, within ~10 s&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Checking the guest side is easy. The balloon device must exist and feature bit 5 (&lt;code>VIRTIO_BALLOON_F_REPORTING&lt;/code>) must be negotiated:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;span class="lnt">2
&lt;/span>&lt;span class="lnt">3
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-bash" data-lang="bash">&lt;span class="line">&lt;span class="cl">&lt;span class="k">for&lt;/span> d in /sys/bus/virtio/devices/*&lt;span class="p">;&lt;/span> &lt;span class="k">do&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="o">[&lt;/span> &lt;span class="s2">&amp;#34;&lt;/span>&lt;span class="k">$(&lt;/span>cat &lt;span class="nv">$d&lt;/span>/device&lt;span class="k">)&lt;/span>&lt;span class="s2">&amp;#34;&lt;/span> &lt;span class="o">=&lt;/span> 0x0005 &lt;span class="o">]&lt;/span> &lt;span class="o">&amp;amp;&amp;amp;&lt;/span> cut -c6 &lt;span class="nv">$d&lt;/span>/features &lt;span class="c1"># 1 = reporting on&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">done&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>If you run Proxmox 8 or 9 with Linux guests, the fix is simply: do not set &lt;code>balloon: 0&lt;/code>. Delete the line from any VM where you want memory to come back. You lose nothing except the ability to stop PVE from ballooning under host pressure, and you can still set a &lt;code>balloon&lt;/code> minimum equal to the VM memory for that.&lt;/p>
&lt;p>Two more things kill it silently:&lt;/p>
&lt;ul>
&lt;li>&lt;code>hugepages:&lt;/code> in the VM config. The host cannot discard sub-hugepage ranges, so reports are pointless.&lt;/li>
&lt;li>Guest kernels without &lt;code>CONFIG_PAGE_REPORTING&lt;/code> or without a free page reporting capable balloon driver. Linux 5.7+ has it. FreeBSD and Windows guests on the same host showed RSS equal to their full memory despite the QEMU flag, because their balloon drivers do not implement it.&lt;/li>
&lt;/ul>
&lt;h2 id="what-free-page-reporting-does-not-cover">What Free Page Reporting Does Not Cover
&lt;/h2>&lt;p>With reporting on, process exit behaves like a container. But two kinds of memory never become &amp;ldquo;free&amp;rdquo; in the guest, and so are never reported.&lt;/p>
&lt;h3 id="page-cache">Page cache
&lt;/h3>&lt;p>Linux keeps file pages cached after the process that read them exits. From the host&amp;rsquo;s point of view those are in-use guest pages. Reading 512 MiB of files inside the guest:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Step&lt;/th>
&lt;th>Guest &lt;code>buff/cache&lt;/code>&lt;/th>
&lt;th>Host RSS&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>idle&lt;/td>
&lt;td>37 MiB&lt;/td>
&lt;td>249 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>after reading 512 MiB&lt;/td>
&lt;td>549 MiB&lt;/td>
&lt;td>753 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>20 s later&lt;/td>
&lt;td>549 MiB&lt;/td>
&lt;td>753 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>echo 1 &amp;gt; /proc/sys/vm/drop_caches&lt;/code>&lt;/td>
&lt;td>21 MiB&lt;/td>
&lt;td>255 MiB&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Normal reclaim only runs under memory pressure. A guest with headroom never has pressure, so the cache sits there forever. This is the real reason idle VMs look fat on the host.&lt;/p>
&lt;h3 id="fragmented-free-memory">Fragmented free memory
&lt;/h3>&lt;p>Reporting works on contiguous blocks of &lt;code>page_reporting_order&lt;/code> pages. The default is the pageblock order, &lt;code>9&lt;/code> on x86, which is 2 MiB. If a workload frees memory in pieces that never form a 2 MiB block, nothing is reported. More on this below, because it turned out to matter a lot.&lt;/p>
&lt;h2 id="how-orbstack-does-it">How OrbStack Does It
&lt;/h2>&lt;p>I looked inside the OrbStack VM to see what they changed. The short version: a patched kernel, a tuned reporting granularity, and their own VMM.&lt;/p>
&lt;p>&lt;strong>The balloon side is the same idea.&lt;/strong> The virtio-balloon negotiates &lt;code>STATS&lt;/code>, &lt;code>FREE_PAGE_HINT&lt;/code> and &lt;code>REPORTING&lt;/code>, and dmesg says &lt;code>Free page reporting enabled&lt;/code>. The host side is OrbStack&amp;rsquo;s own Rust device model (the helper binary contains &lt;code>src/devices/src/virtio/balloon/device.rs&lt;/code> and strings like &lt;code>free-page reporting queue event&lt;/code>), so they handle the reports themselves rather than relying on Apple&amp;rsquo;s framework devices.&lt;/p>
&lt;p>&lt;strong>Reporting granularity is 16 KiB, not 2 MiB.&lt;/strong>&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;span class="lnt">2
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-fallback" data-lang="fallback">&lt;span class="line">&lt;span class="cl">$ cat /sys/module/page_reporting/parameters/page_reporting_order
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">2
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>Order 2 on a 4 KiB page kernel is 16 KiB, which is exactly the macOS page size. Fragmentation stops being a problem because almost every freed page lands in a reportable block. They also set &lt;code>compaction_proactiveness=0&lt;/code>, which they can afford because they no longer need to form 2 MiB blocks.&lt;/p>
&lt;p>&lt;strong>A kernel thread reclaims page cache on a timer.&lt;/strong> This is the part a stock kernel does not have. There is a kernel thread literally called &lt;code>reclaim&lt;/code> (pid 130). It is hidden from &lt;code>/proc&lt;/code> along with every other kernel thread, but it shows up in &lt;code>/sys/kernel/debug/sched/debug&lt;/code>:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;span class="lnt">2
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-fallback" data-lang="fallback">&lt;span class="line">&lt;span class="cl"> S reclaim 130 ... 40 switches, 569 ms total runtime
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> S kswapd0 71 ... 3 switches, 14 ms total runtime
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>I enabled the &lt;code>vmscan&lt;/code> tracepoints and read several GiB of files for 75 seconds. Every single one of the 1100 &lt;code>mm_vmscan_lru_shrink_inactive&lt;/code> events came from task &lt;code>reclaim&lt;/code>. The &lt;code>pgsteal_kswapd&lt;/code>, &lt;code>pgsteal_direct&lt;/code> and &lt;code>pgsteal_proactive&lt;/code> counters stayed at zero the whole time. The thread fires roughly every 45 seconds and holds page cache at about 100 to 250 MiB no matter what. Sampling &lt;code>pgsteal_file&lt;/code> once a second:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;span class="lnt">2
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-fallback" data-lang="fallback">&lt;span class="line">&lt;span class="cl">t=40s +15263 pages cached=162MiB
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">t=85s +22118 pages cached=80MiB
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>Supporting settings: MGLRU enabled with &lt;code>min_ttl_ms=500&lt;/code>, &lt;code>swappiness=20&lt;/code>, THP in &lt;code>madvise&lt;/code> mode, a 16 GiB zram swap plus a 1 GiB disk swap.&lt;/p>
&lt;p>&lt;strong>What it does not do.&lt;/strong> Two things, for honesty&amp;rsquo;s sake:&lt;/p>
&lt;ul>
&lt;li>Idle anonymous memory is not reclaimed. A container holding 1.5 GiB untouched for 4 minutes never went to zram. Only freed memory and page cache go back.&lt;/li>
&lt;li>Reclaimable slab is not trimmed. A &lt;code>find&lt;/code> over the Docker storage left 3.6 GiB of dentry and inode cache that sat there for over 20 minutes, with the macOS footprint stuck at 5.9 GiB. &lt;code>echo 2 &amp;gt; /proc/sys/vm/drop_caches&lt;/code> released it to macOS within 10 seconds.&lt;/li>
&lt;/ul>
&lt;p>So &amp;ldquo;immediate give-back&amp;rdquo; is specifically about process exit and file cache. That is also what matters most in practice.&lt;/p>
&lt;h2 id="reproducing-the-pieces-on-proxmox-ve">Reproducing the Pieces on Proxmox VE
&lt;/h2>&lt;h3 id="lower-the-reporting-order">Lower the reporting order
&lt;/h3>&lt;p>The &lt;code>page_reporting_order&lt;/code> parameter exists in stock kernels and is writable at runtime, no reboot:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-bash" data-lang="bash">&lt;span class="line">&lt;span class="cl">&lt;span class="nb">echo&lt;/span> &lt;span class="m">2&lt;/span> &amp;gt; /sys/module/page_reporting/parameters/page_reporting_order
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>To persist it, add &lt;code>page_reporting.page_reporting_order=2&lt;/code> to the guest kernel command line.&lt;/p>
&lt;p>Does it matter? I wrote a small test: allocate 1200 MiB of private anonymous memory, then free every other 1 MiB slice with &lt;code>madvise(MADV_DONTNEED)&lt;/code>. That leaves 600 MiB free in order-8 blocks which can never coalesce into order-9. After 25 seconds the process frees everything.&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt"> 1
&lt;/span>&lt;span class="lnt"> 2
&lt;/span>&lt;span class="lnt"> 3
&lt;/span>&lt;span class="lnt"> 4
&lt;/span>&lt;span class="lnt"> 5
&lt;/span>&lt;span class="lnt"> 6
&lt;/span>&lt;span class="lnt"> 7
&lt;/span>&lt;span class="lnt"> 8
&lt;/span>&lt;span class="lnt"> 9
&lt;/span>&lt;span class="lnt">10
&lt;/span>&lt;span class="lnt">11
&lt;/span>&lt;span class="lnt">12
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-python" data-lang="python">&lt;span class="line">&lt;span class="cl">&lt;span class="kn">import&lt;/span> &lt;span class="nn">mmap&lt;/span>&lt;span class="o">,&lt;/span> &lt;span class="nn">time&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="n">MB&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="mi">1024&lt;/span> &lt;span class="o">*&lt;/span> &lt;span class="mi">1024&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="n">n&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="mi">1200&lt;/span> &lt;span class="o">*&lt;/span> &lt;span class="n">MB&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="n">m&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="n">mmap&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">mmap&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="o">-&lt;/span>&lt;span class="mi">1&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">n&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">flags&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="n">mmap&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">MAP_PRIVATE&lt;/span> &lt;span class="o">|&lt;/span> &lt;span class="n">mmap&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">MAP_ANONYMOUS&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">for&lt;/span> &lt;span class="n">i&lt;/span> &lt;span class="ow">in&lt;/span> &lt;span class="nb">range&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="mi">0&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">n&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="mi">4096&lt;/span>&lt;span class="p">):&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">m&lt;/span>&lt;span class="p">[&lt;/span>&lt;span class="n">i&lt;/span>&lt;span class="p">]&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="mi">1&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="nb">print&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="s2">&amp;#34;allocated&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">flush&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="kc">True&lt;/span>&lt;span class="p">);&lt;/span> &lt;span class="n">time&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">sleep&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="mi">12&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">for&lt;/span> &lt;span class="n">off&lt;/span> &lt;span class="ow">in&lt;/span> &lt;span class="nb">range&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="mi">0&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">n&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="mi">2&lt;/span> &lt;span class="o">*&lt;/span> &lt;span class="n">MB&lt;/span>&lt;span class="p">):&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">m&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">madvise&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="n">mmap&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">MADV_DONTNEED&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">off&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">MB&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="nb">print&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="s2">&amp;#34;freed 600M in 1MiB slices&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">flush&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="kc">True&lt;/span>&lt;span class="p">);&lt;/span> &lt;span class="n">time&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">sleep&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="mi">25&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="n">m&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">close&lt;/span>&lt;span class="p">()&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="nb">print&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="s2">&amp;#34;freed rest&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">flush&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="kc">True&lt;/span>&lt;span class="p">);&lt;/span> &lt;span class="n">time&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">sleep&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="mi">20&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Reporting order&lt;/th>
&lt;th>Host RSS after the fragmented 600 MiB free&lt;/th>
&lt;th>after full free&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>9 (default, 2 MiB)&lt;/td>
&lt;td>1434 MiB, &lt;strong>unchanged&lt;/strong>&lt;/td>
&lt;td>282 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>2 (16 KiB)&lt;/td>
&lt;td>&lt;strong>853 MiB&lt;/strong>, within ~12 s&lt;/td>
&lt;td>243 MiB&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>With the default order, the fragmented 600 MiB is invisible to the host until the whole process exits. With order 2 nearly all of it comes back. The cost is more reporting traffic and more &lt;code>madvise&lt;/code> calls on the host.&lt;/p>
&lt;h3 id="watch-out-for-transparent-huge-pages">Watch out for Transparent Huge Pages
&lt;/h3>&lt;p>My first two runs of that test showed nothing freed at all, and the reason is worth its own warning. The guest had THP set to &lt;code>always&lt;/code>. When a process frees part of a 2 MiB huge page, the kernel splits the page table mapping and drops it from the process&amp;rsquo;s &lt;code>AnonPages&lt;/code>, but the physical huge page is parked on a deferred-split queue until memory pressure runs the shrinker. &lt;code>AnonPages&lt;/code> dropped by 200 MiB while &lt;code>MemFree&lt;/code> did not move by a single megabyte:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-fallback" data-lang="fallback">&lt;span class="line">&lt;span class="cl">anon: 7 -&amp;gt; touched 407 -&amp;gt; after madvise 207 MiB; free: 1844 -&amp;gt; 1456 -&amp;gt; 1456
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>From the host&amp;rsquo;s point of view nothing was freed, regardless of reporting order. Any workload that frees memory in pieces smaller than 2 MiB hits this, which is most of them. OrbStack runs &lt;code>madvise&lt;/code>. On guests where give-back matters, I would do the same:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-bash" data-lang="bash">&lt;span class="line">&lt;span class="cl">&lt;span class="nb">echo&lt;/span> madvise &amp;gt; /sys/kernel/mm/transparent_hugepage/enabled
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;h3 id="trim-page-cache-without-a-kernel-patch">Trim page cache without a kernel patch
&lt;/h3>&lt;p>The OrbStack &lt;code>reclaim&lt;/code> thread is a kernel patch I cannot copy, but there are three stock ways to get something similar. All of them are LRU-based, so they evict least recently used pages first, unlike &lt;code>drop_caches&lt;/code> which throws everything away.&lt;/p>
&lt;p>&lt;strong>&lt;code>memory.reclaim&lt;/code> on a timer.&lt;/strong> cgroup v2 kernels from 5.19 expose &lt;code>memory.reclaim&lt;/code>. Writing an amount runs normal reclaim for that much. It works on the root cgroup, and &lt;code>swappiness=0&lt;/code> restricts it to file pages (6.6+):&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-bash" data-lang="bash">&lt;span class="line">&lt;span class="cl">&lt;span class="nb">echo&lt;/span> &lt;span class="s2">&amp;#34;256M swappiness=0&amp;#34;&lt;/span> &amp;gt; /sys/fs/cgroup/memory.reclaim
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>Tested on the Alpine guest: cache went from 137 to 97 MiB with a 32 MiB request, nothing else disturbed. A cron job doing this every minute is the cheap OrbStack imitation.&lt;/p>
&lt;p>&lt;strong>&lt;code>memory.high&lt;/code> on a workload cgroup.&lt;/strong> Put your containers or services in a cgroup with a &lt;code>memory.high&lt;/code> ceiling. Reclaim then runs at the limit continuously, and freed cache becomes reportable. The downside is you have to guess the limit.&lt;/p>
&lt;p>&lt;strong>DAMON reclaim.&lt;/strong> The most elegant, because it is time-based rather than amount-based or limit-based. DAMON samples one page per region to find memory that has not been accessed for &lt;code>min_age&lt;/code> (default 2 minutes) and pages it out, with bounded CPU cost. That is container-like semantics: memory footprint tracks the actual working set with a lag you choose. Two details matter in a VM. &lt;code>damon_reclaim&lt;/code> ships with watermarks that keep it idle unless free memory is between 20% and 50% of RAM, so on a guest with headroom (the case we care about) it does nothing until &lt;code>wmarks_high&lt;/code> and &lt;code>wmarks_mid&lt;/code> are raised to 1000. And it pages out through the ordinary reclaim path, so cold anonymous memory needs swap to go anywhere: with zram in the guest it ends up compressed at 3–4:1 and only the saved part is reported to the host, without swap only page cache can be reclaimed. The catch is kernel support. From the configs I checked:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Kernel&lt;/th>
&lt;th>&lt;code>CONFIG_DAMON_RECLAIM&lt;/code>&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Debian 13 trixie 6.12&lt;/td>
&lt;td>yes&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Fedora&lt;/td>
&lt;td>yes&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Ubuntu 24.04 generic 6.8&lt;/td>
&lt;td>no&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Proxmox VE kernels (6.17, 7.0)&lt;/td>
&lt;td>no&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Alpine lts / virt&lt;/td>
&lt;td>no&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Raspberry Pi OS 6.6&lt;/td>
&lt;td>no&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>So for a Debian guest, &lt;code>damon_reclaim&lt;/code> is a module parameter away. Everywhere else, use &lt;code>memory.reclaim&lt;/code>, or rebuild the kernel: for Alpine that turned out to be a Dockerfile and a six-line config fragment, described in the &lt;a class="link" href="https://charlie0129.github.io/blog/p/alpine-image-builder/#a-kernel-of-your-own" >image builder post&lt;/a>.&lt;/p>
&lt;h2 id="bonus-moving-the-page-cache-to-the-host-with-virtio-pmem">Bonus: Moving the Page Cache to the Host with virtio-pmem
&lt;/h2>&lt;p>This started as a question about Kata Containers, which use virtio-fs with DAX so that guests do not keep their own page cache. Plain virtio-fs does not do that, and the DAX patches never landed in upstream QEMU. But QEMU does ship &lt;code>virtio-pmem&lt;/code>, which is the same idea applied to a disk image: a host file mapped into guest physical memory. The guest mounts ext4 with &lt;code>-o dax&lt;/code>, and file data is served straight from the host page cache with no guest copy.&lt;/p>
&lt;p>Proxmox has no config knob for it, so it is an &lt;code>args:&lt;/code> line. The VM also needs &lt;code>hotplug: memory&lt;/code> and &lt;code>numa: 1&lt;/code> so a device-memory region exists:&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt">1
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-gdscript3" data-lang="gdscript3">&lt;span class="line">&lt;span class="cl">&lt;span class="n">args&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="o">-&lt;/span>&lt;span class="n">object&lt;/span> &lt;span class="n">memory&lt;/span>&lt;span class="o">-&lt;/span>&lt;span class="n">backend&lt;/span>&lt;span class="o">-&lt;/span>&lt;span class="n">file&lt;/span>&lt;span class="p">,&lt;/span>&lt;span class="n">id&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="n">pmem0&lt;/span>&lt;span class="p">,&lt;/span>&lt;span class="n">share&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="n">on&lt;/span>&lt;span class="p">,&lt;/span>&lt;span class="n">mem&lt;/span>&lt;span class="o">-&lt;/span>&lt;span class="n">path&lt;/span>&lt;span class="o">=/&lt;/span>&lt;span class="k">var&lt;/span>&lt;span class="o">/&lt;/span>&lt;span class="n">lib&lt;/span>&lt;span class="o">/&lt;/span>&lt;span class="n">vz&lt;/span>&lt;span class="o">/&lt;/span>&lt;span class="n">images&lt;/span>&lt;span class="o">/&lt;/span>&lt;span class="mi">300&lt;/span>&lt;span class="o">/&lt;/span>&lt;span class="n">vm&lt;/span>&lt;span class="o">-&lt;/span>&lt;span class="mi">300&lt;/span>&lt;span class="o">-&lt;/span>&lt;span class="n">pmem0&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">raw&lt;/span>&lt;span class="p">,&lt;/span>&lt;span class="n">size&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="mi">4&lt;/span>&lt;span class="n">G&lt;/span> &lt;span class="o">-&lt;/span>&lt;span class="n">device&lt;/span> &lt;span class="n">virtio&lt;/span>&lt;span class="o">-&lt;/span>&lt;span class="n">pmem&lt;/span>&lt;span class="o">-&lt;/span>&lt;span class="n">pci&lt;/span>&lt;span class="p">,&lt;/span>&lt;span class="n">memdev&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="n">pmem0&lt;/span>&lt;span class="p">,&lt;/span>&lt;span class="n">id&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="n">nv0&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>The guest kernel needs &lt;code>CONFIG_FS_DAX&lt;/code>. Alpine&amp;rsquo;s &lt;code>linux-virt&lt;/code> does not have it, &lt;code>linux-lts&lt;/code> does. I cloned the root onto &lt;code>/dev/pmem0&lt;/code> and booted with &lt;code>root=/dev/pmem0 rootfstype=ext4 rootflags=dax&lt;/code>. Results:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Step&lt;/th>
&lt;th>Guest page cache&lt;/th>
&lt;th>Host RssAnon&lt;/th>
&lt;th>Host RssFile&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>fresh boot, 740 MiB of root files&lt;/td>
&lt;td>25 MiB&lt;/td>
&lt;td>355 MiB&lt;/td>
&lt;td>32 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>after reading every file on &lt;code>/&lt;/code>&lt;/td>
&lt;td>48 MiB&lt;/td>
&lt;td>371 MiB&lt;/td>
&lt;td>724 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>host cgroup &lt;code>memory.high=550M&lt;/code> for 20 s&lt;/td>
&lt;td>48 MiB&lt;/td>
&lt;td>94 MiB&lt;/td>
&lt;td>299 MiB&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Reading 700 MiB of files added 24 MiB of guest cache. The data lives in the host page cache of the backing file instead, which is ordinary file-backed memory: when I squeezed the VM&amp;rsquo;s cgroup, the host evicted it and the guest did not notice. Podman with kernel overlayfs on top worked without a single warning (the overlayfs trouble people remember is specific to virtio-fs, where the upper layer is FUSE). The backing store can also be an LVM volume, I tested that too, since &lt;code>mem-path&lt;/code> only needs something mmap-able.&lt;/p>
&lt;p>Caveats: no live migration, no vzdump backup of that disk, the guest warns it cannot guarantee write persistence because durability depends on fsync reaching the host file, and judging the VM&amp;rsquo;s memory cost now requires looking at &lt;code>RssAnon&lt;/code> rather than &lt;code>VmRSS&lt;/code>. It is a fun experiment and genuinely effective, but for a normal PVE VM I would stop at the previous section.&lt;/p>
&lt;h2 id="summary">Summary
&lt;/h2>&lt;p>If you want VMs on Proxmox VE to return memory like containers:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Never set &lt;code>balloon: 0&lt;/code>&lt;/strong> on Linux guests. That one line disables free page reporting, which PVE otherwise enables for you.&lt;/li>
&lt;li>&lt;strong>Set &lt;code>page_reporting.page_reporting_order=2&lt;/code>&lt;/strong> in the guest if your workloads free memory in small pieces. Measured: 600 MiB that was invisible to the host came back within 12 seconds.&lt;/li>
&lt;li>&lt;strong>Set THP to &lt;code>madvise&lt;/code>&lt;/strong> in the guest. With &lt;code>always&lt;/code>, partial frees inside huge pages are not freed at all until memory pressure.&lt;/li>
&lt;li>&lt;strong>Trim page cache proactively&lt;/strong>: &lt;code>damon_reclaim&lt;/code> on Debian, or &lt;code>memory.reclaim&lt;/code> on a timer elsewhere.&lt;/li>
&lt;li>Do not use &lt;code>hugepages:&lt;/code> on those VMs, and accept that FreeBSD and Windows guests will not participate.&lt;/li>
&lt;/ol>
&lt;p>That gets you within a few hundred MiB of a VM&amp;rsquo;s real working set, with memory returning seconds after a process exits.&lt;/p>
&lt;p>For Alpine guests, points 2 to 4 are now a hook in my &lt;a class="link" href="https://charlie0129.github.io/blog/p/alpine-image-builder/#hooks" >image builder&lt;/a>: add &lt;code>66-vmmem&lt;/code> to &lt;code>HOOKS&lt;/code> and the image boots with the reporting order at 2, THP in &lt;code>madvise&lt;/code> mode, and a one-minute &lt;code>memory.reclaim&lt;/code> job that keeps the page cache at a configurable floor. With the builder&amp;rsquo;s DAMON-enabled kernel, &lt;code>VMMEM_DAMON_RECLAIM=yes&lt;/code> adds point 4&amp;rsquo;s time-based reclaim as well.&lt;/p></description></item></channel></rss>