The usual explanation is that mipmaps blur things as they get further away. That belongs to the wrong family of ideas.
A mipmap actually presents in the scene as the legal band-limit of whatever albedo the pixel grid can still hold. Skip the mip chain, and high-frequency details—wood grain, brick, small glyphs, and tiles—continue feeding the shader frequencies that a single sample per pixel simply cannot represent. Consequently, those frequencies fold. Distant floors begin to sparkle, crawl, and moiré, while specular highlights aggressively glitter. If you put a box pyramid under a trilinear sampler, those false patterns collapse into soft, legally represented wood, brick, and disks. Crucially, the near field—where magnification happens—must still look identical.
The Cornell box and the hallway photographs demonstrate this principle in a physical space. Afterward, the zone-plate FFT proves exactly why the missing energy was folded rather than artistically blurred.
Look at the Cornell box rendered from the exact same camera angle.

On the left, using GL_NEAREST with no mipmap, the satin wood planks, letter-grid wallpaper, and brushed metal are all authored finer than the distant pixels can hold. As a result, the floor sparkles, the small glyphs degrade into moiré, and the metal glitters. On the right, we apply a CPU box pyramid to both the albedo and roughness maps using GL_LINEAR_MIPMAP_LINEAR. Those false high-frequency patterns collapse. What remains is the actual wood texture the pixel grid can legally represent. Both the tall and short boxes are close enough that their sides still read clearly. This is exactly how a mipmap presents. It is not a fog effect, nor is it a simple “far therefore blur” heuristic; it is a pre-band-limited level of detail.
The everyday extension of this concept is walking down a familiar hallway.

The folding artifact is most obvious on the vanishing floor. Notice that the near door frames still match perfectly between the two shots—because magnification is purely reconstruction. The leftover softness on the right side is simply isotropic over-blur, where band-limits the texture to its major axis. This makes no claims about anisotropic filtering (AF) or how NVIDIA explicitly selects a mip level.
A distance strip effectively separates these two regimes, which is highly visible in the wood grain:

The NEAR column must match. If the “mip” version looks softer on a close floor, the implementation is broken. The FAR column is where the fold happens: the no-mip grain still crawls, whereas the mip version resolves into the legal plank.
At grazing angles, the Cook-Torrance GGX specular streak on the varnish should still read convincingly as a floor. The fold primarily affects the underlying wood albedo.

Moving in close on the metal box reveals the brush strokes versus the mipmap, the green-wall bounce on the flank, and the floor seams. Because it is in the near-field, the “MIP” decal stays perfectly readable.

Shading relies on Cook-Torrance GGX so the room feels like a physical photograph rather than a sterile Lambertian lab. Lighting consists of point lights, an analytic ceiling rectangle, and a cheap IBL lobe—not a path tracer. It uses linear lighting with an ACES display transform. A CPU box SSAA is applied so the silhouettes read cleanly. None of these specific frames enter the CSV, nor do they receive an FFT. The HUD explicitly reads: PBR photo-only … GGX … rho=n/a. The hero numbers remain locked to the chirp tests.
That covers the visual presentation. Here is the proof that the missing energy was folded.
The lab version of the same sparkle
Consider a circular zone-plate (a Fresnel chirp). Its instantaneous frequency rises with the radius, giving us a known Nyquist ring. Using the exact same texture and orthographic pose, we minify the image at a locked footprint of (): eight texels per pixel, but only one sample per pixel.

The resulting disk is a pure moiré pattern. That is not simply “too far” away. It is high-frequency energy from above the new Nyquist limit folding down and sitting at a lower frequency. It is the exact same phenomenon seen on the Cornell floor, but here it has a known instantaneous frequency .
Using bilinear filtering without a mip chain does not fix the issue. GL_LINEAR is merely a triangle filter. While it takes some of the heat out, it ultimately leaves the fold intact.

Now, let's examine the same footprint using a CPU box pyramid with GL_LINEAR_MIPMAP_LINEAR. For comparison, the right side shows a pyramid that operates strictly as a pyramid without low-passing.

On the left, we see the legal disk—the lab equivalent of the soft wood. On the right, we simply see a smaller picture containing the exact same aliases. A mip chain without a low-pass filter is not a mipmap. The right column is included specifically to ensure that “we generated mip levels” is never confused with “we band-limited the signal.” If we rendered the Cornell box with a point-sampled pyramid, the floor would still sparkle.
Two regimes, never mixed
Magnification (, ) is strictly about reconstruction. Mip and non-mip versions must match perfectly. If “mip” looks softer here, the note is broken.
Minification (, ) is where the band-limit or the fold occurs. This is the core subject of the article.
On this run, at : LINEAR versus LINEAR_MIPMAP_LINEAR yields a mean absolute error (MAE) of . Both samplers are reading from L0 with linear magnification. That column serves as our control, not the hero metric. It acts as the same control mechanism as the NEAR column in the distance strip.

At , the authored chirp has only just started to kiss the new Nyquist limit; still matches closely. However, from onward, the top row degenerates into a lattice of aliases, while the bottom row successfully shrinks into a legal disk. Notably, the pixel MAE during minification is not the defining theorem. By , the MAE is only simply because the no-mip variant has already collapsed into a gray field of aliases. The underlying structures still wildly diverge.
, , and the fold
This lab’s is the GL footprint: texels per pixel. On the ortho science quad it is arithmetic, not dFdx:
Locked . Hero: , , . cannot fit a 1024-wide UV window on a 512 framebuffer; the mag control shows the center texels ().
A texture frequency (cycles/texel) appears on screen at . Screen Nyquist is cycle/px. Fold when
At , anything above cycle/texel is past the new Nyquist. Wood grain, brick edges, glyph stems, tile grout all live in that regime once the surface is small enough on screen. They have nowhere legal to go except a lower . That lower is the sparkle.
The zone-plate is a Fresnel chirp, disk-masked, authored at pixel centers:
Most of that disk has nowhere legal to go, at , except a lower . L0 is authored at or under Nyquist and then measured — not a brick-wall bandlimited field.


A 1-D chirp is the cousin with a single ridge. Phase is locked so ; the naive would have aliased the right half of L0 and we would have been measuring our own authoring bug.


The theorem is the spectrum
The PNG is a 16-entry heat LUT of , where . The actual science relies on calling glReadPixels(..., GL_FLOAT) from an RGBA32F FBO, taking an interior crop, mean-subtracting it, applying a separable Hann window, and executing an unnormalized radix-2 DFT within this binary. Power is calculated as . We use the exact same vmin/vmax scaling on the pair so they can be accurately compared. Do not attempt to FFT the Cornell PNG yourself.
Using nearest-neighbor with no mipmap results in a complete fold:

Using bilinear filtering with no mipmap still runs hot at high , retaining a vmax of :

Comparing the box-mip against a point-subsample, using the exact same scale as 05:

That specific image pair visually represents the theorem. The high- energy clearly present in 05 is completely missing from 09-left. Meanwhile, 09-right is identical to 05, because a point-sampled pyramid never actually passes through a low-pass filter. The Cornell floor became soft due to the mathematics in 09-left, not 09-right.
A box filter in spatial coordinates equates to a sinc filter in the frequency domain:
The resulting sidelobes above the new Nyquist limit are distinctly visible in 09-left. That artifact is a mark of honesty, not a driver bug. We did not use glGenerateMipmap; Mesa's internal path functions as a blit/box-ish operation anyway. Both the semantic albedo and ORM maps rely on this same CPU box. The point-subsample acts purely as the control:
Trilinear filtering is defined as . This note makes no specific claims about commercial hardware filtering quality.
Quote . Do not quote as .
Radial
with . Because annuli are weighted equally, the Hann/box-sinc lobes dwelling in the outer rings prevent the ratio from collapsing, even after the fold is eliminated.
At , , using a crop of 256, and padded=1, we measured the following on this run:
| filter | vmax | ||
|---|---|---|---|
NEAREST no-mip |
0.274 | 898 | 2.21 |
LINEAR no-mip |
0.256 | 379 | 2.21 |
| box-mip trilinear | 0.241 | 30 | 0.998 |
| point-subsample | 0.328 | 899 | 2.22 |
textureLod |
0.241 | 30 | 0.998 |
Notice that only moves from to . That is emphatically not a drop. The pixel-weighted similarly barely moves from to . At , the radial is actually higher for the box-mip than for LINEAR ( vs ), even though plummets from to .
represents the sum of excluding DC. 898 vs 30 represents an approximate reduction. That significant drop, alongside the shared-scale pair, forms the core technical claim. The textureLod output perfectly matched the implicit-LOD photograph on this orthographic build. Do not attribute these exact metrics to the wood floor visuals; the Cornell box serves as the presentation, whereas the chirp acts as the scientific meter.

When the quad is smaller than the crop window ( or ), the framebuffer clears to —the mean of the zone-plate. Because of this, the padding resolves to zeros after mean-subtraction (padded=1 in the CSV). Using a hard rectangular cut would have introduced a 2-D sinc that dominates every filter in the test.
The Hann window applied to every crop possesses its own spectrum, preventing its cross pattern from being mistaken for genuine aliasing:

Spectrum follows , not “distance”
There is no camera operating on the orthographic path. By adjusting the lodBias by at a fixed footprint:

| bias | ||
|---|---|---|
| 2 | 110 | |
| 3 | 30 | |
| 4 | 0.27 |
Under-biasing clearly leaves a residual fold. Over-biasing strips away almost all AC power. The exact same cells were verified using textureLod. This demonstrates a direct level pick. The softness observed in the Cornell box is driven by the exact same knob: the specific prefiltered level the sampler lands on dictates the blur, not the raw metric distance the camera moved.
If we force the mip level on a window, the disk physically shrinks because demands it, not because the quad shifted in space:

The values along that strip are: 1146, 1042, 815, 437, and 144. The residual box-sinc lobes visible at represent the same honesty gap shown earlier in 09-left.
Written in GLSL 330 using a single uniform:
vec4 s = (uMode == 1) ? textureLod(uTex, vUV, uLod)
: texture(uTex, vUV, uBias);
textureQueryLod is out (GLSL 400). The LOD photograph is composed of a falsecolor mip chain combined with NEAREST_MIPMAP_NEAREST. Running on llvmpipe with GALLIVM_PERF=no_filter_hacks, the implicit directly matched the CPU ladder. Do not take that to NVIDIA / AMD / Intel. Furthermore, do not assume dFdx perfectly aligns with hardware derivatives. The scientific path simply does not care: it relies on CPU , a CPU pyramid, and explicit textureLod.

Isotropic mip is the wrong ellipse
One photograph, no on the HUD, no AF table. This perfectly illustrates the residual softness found on the hallway's mip side, visually explained via a zone-plate.

Because strictly band-limits to the major axis, it heavily over-blurs the minor axis. EWA and anisotropic filtering (AF) are the more advanced cousins that this specific lab note does not run. Keep in mind that llvmpipe is not an appropriate backend for a discrete-GPU quality study.
What the PBR path actually is
The shading pipeline uses Cook-Torrance GGX with a metalness-roughness workflow in GLSL 330. Specifically, is GGX / Trowbridge-Reitz, is Smith with Schlick-GGX, and is Schlick. It relies entirely on direct lighting—it is not a path tracer. There is no Toksvig mapping. Both the albedo and ORM maps receive the exact same CPU box filter. The display pipeline applies ACES (Narkowicz) followed by a curve from a linear FBO.
The room photographs were captured at a high gallery resolution of utilizing CPU box SSAA, rather than GL MSAA. This SSAA pass is dedicated solely to geometry—smoothing out box edges, the lamp quad, door frames, and hallway vanishing lines. While it slightly averages out the no-mip sparkle, the core visual difference between the fold and the band-limit remains distinctly clear. The zone-plate and FFT frames are locked to with a FBO.
The Cornell setup provides a familiar structural layout; it is not intended as a global illumination benchmark. The combination of an area light, bounce points, and gradient IBL strictly establishes the look. Ultimately, rendering in PBR does not alter the underlying math of the theorem.
What this box actually measured
Host: OSMesa, Mesa 25.0.7-2+deb13u1, llvmpipe (LLVM 19.1.7, 256 bits). The FBO uses an RGBA32F format, meaning the 8-bit fallback is not hit. There is no sRGB and no MSAA applied to the science FBO. Texture wrap is set to CLAMP_TO_EDGE. Out of all internal assertions, 27 pass / 0 fail. This includes the DFT self-test, verifying MAE , and ensuring the for nearest is that of the mipmap (898 vs 30). The semantic scene frames are purely supplementary outputs; they are not part of the core scientific checks.
We can claim: On this specific rasterizer, minifying an authored chirp without a mip chain forces energy to fold. Applying a CPU box pyramid alongside trilinear sampling removes the vast majority of that AC power. We measured this directly using the binary's DFT calculated from a float readback. Furthermore, the sampler pictures—including the PBR Cornell box and the hallway—are accurate representations of this rasterizer's output.
We cannot claim: hardware LOD, anisotropy quality, texel cache performance, bandwidth, occupancy, or definitively state "this is how commercial GPUs work." We cannot assert that the zone-plate acts as a perfect LPF source. We cannot simply FFT the resulting PNG and label it as rigorous science. Finally, we cannot directly attribute the numerical results to the visual softness of the wood floor.
A box filter an ideal LPF. A point-subsample a valid mipmap. the drop. If you do not pad with the mean, the windowing artifact will dominate the results. Magnification MAE must inherently be zero, or the entire methodology is broken. Ultimately, a mipmap presents in the scene as a legal band-limit, never just as distance-fog.