How many samples does this shot need? Nobody measures. They guess.
The advice everyone gives — including the documentation of the tools thatsell you faster renders — is this: render, drop the samples, look at it, repeatuntil it goes bad. That is the most common question in Cycles, answered bysquinting.
SampleProof answers it with a measurement, on your scene, against areference it renders itself.
First, the uncomfortable part: there is no magic number
This add-on started out looking for the point where a frame stops changing.It does not exist, and the measurement is what proved it.
Path-traced noise falls as one over the square root of the samples. Thatmeans the difference between your render and a converged one keeps shrinkingforever, by the same fraction every time you double, with no knee and noplateau. Measured on two scenes — a plain one and one made of glass — the ratiobetween neighbouring steps came out 1.42, 1.55, 1.62, 1.61 and 1.27, 1.41,1.28, 1.38. The theory says 1.41. Both scenes behave exactly like the theory,and neither has a point where more samples stop helping.
So a tool that hands you a sample count without asking what you will accepthas quietly made your decision for you. This one refuses to do that, and givesyou the three things that actually let you decide.
1. Where you are right now
At 64 samples you are 0.36 % from converged. That is a plain scene.The same 64 samples on a glass-and-point-light scene read 1.35 % — nearly fourtimes further. Same number in the box, completely different picture, and untilyou measure you have no way to know which one you are in.
2. What one more doubling buys on this shot
Doubling the samples cuts the remaining difference by 35 %. Measuredacross your own ladder, not assumed from theory — if your scene behavesdifferently, you see that. This is the exchange rate: every doubling costs twicethe render time and buys the same fraction. Now the decision is a decision, nota guess.
3. The number that meets your tolerance
You say how close to converged is close enough. At 0.5 % the plain scene issatisfied by 64 samples. The glass scene is not satisfied by 64 at all, andSampleProof says so and says what it would take: about four more doublings,roughly 1024 samples. It does not print a number it cannot stand behind.
And where in the frame the difference lives
The frame is divided into a grid and the noisiest cell is named, along withthe object the camera sees through it. If one glass prop is holding an entireshot hostage, fixing that prop is cheaper than raising samples for every pixelin the frame.
What the denoiser really costs
One extra frame is rendered with denoising on and compared against theconverged reference, not against the noisy frame. That is thecomparison that matters: it tells you what the denoiser saves in noiseand what it costs in real detail. A denoised frame that reads furtherfrom the reference has smoothed away something that was actually there.
Measured cheaply, and honestly about how
The ladder is rendered at a reduced size, because noise per pixel does notdepend on resolution — a smaller frame carries the same per-pixel noise andpredicts the full one. Measured: 1.065 % against 1.023 % on a full-size frame,seventeen times faster.
Where that breaks is measured too, and said out loud rather than buried: on ashot made of thin wires, reducing 1200 px to 240 overstated the difference by 74to 79 per cent, because detail finer than a pixel does not survive thereduction. So the measurement size is a control in the panel, not a constant.On caustics the same reduction is accurate to within one per cent.
A crop would have been the obvious shortcut, and it is wrong. The same framemeasured on a central crop read 7.973 % — off by seven percentage points,because the crop landed on the glass and reported the glass instead of the shot.Both routes were measured before one was chosen.
It refuses rather than guesses
- Not Cycles. Sample convergence in this sense is a Cyclesidea, and the add-on says so instead of producing a number.
- No camera. Nothing to render, nothing to converge.
- Animated Seed on. Every step would use a different noisepattern and the ladder would not be comparable with itself.
- A reference no cleaner than the ladder. Then there isnothing to compare against, and it says that rather than dressing noise up asan answer.
- Adaptive Sampling. Not a refusal — but it is switched offfor the measurement, named in the report along with your noise threshold, andput back. With it on, the sample count is only a ceiling and the thresholddecides: the top of the ladder flattens to 0.823 % against 0.817 % where ahealthy ladder reads 0.606 against 0.381.
- A measurement that would cost more than the render. Thereference is eight times the top of your ladder, twice — at a top of 4096 thatis over sixty thousand samples. It is refused before the first frame, with thenumber named.
- A compositor with no output node. Blender renders the frameand writes nothing. That used to be a crash; now it is a sentence explainingwhat happened.
- Pixels that are not numbers. One nan from a broken materialused to turn the whole answer into nan. They are left out and counted, and theirpresence is reported.
The reference's own noise is taken back out
Every comparison contains the noise of both frames, so the reference's shareis measured — by rendering it twice with different seeds — and subtracted. Araw 0.3810 % against a reference whose own noise is 0.1186 % becomes 0.3717 %,and that matches an independent calculation to the fourth decimal. The numbersyou read are the frame's, not the measurement's.
It measures the frame you actually render
Your denoiser setting, your render border, your compositor. The ladder runswith denoising exactly as your scene has it, because the answer has to be aboutthe frame you will produce — measured, that changes the verdict from "no samplecount is enough" to "sixteen will do". A render border is kept and named, notquietly removed: on a glass scene, measuring the whole frame instead of theregion being rendered reported 1.869 % where the truth inside the border was0.291 %. A compositor is named too, because what gets measured then is thecomposite and a Denoise node in there changes the answer.
It puts your settings back
It has to render, so it has to touch your render settings — and every one ofthem is recorded before and restored after: samples, seed, resolution, border,output path, file format, denoising. The panel prints the result of thatcomparison every time, not only when something went wrong. The single thing itchanges on purpose is the sample count, and only when you press Use ThatNumber.
Tested
98 checks on each of Blender 4.2.9, 5.1 and 5.2, zero problems, plus twentydeliberate sabotages of the product itself to prove the tests can fail — all ofthem turn the suite red.
An adversarial review of this add-on found four things that would haveshipped as confident lies, and every one is now closed with a measurement and aregression test. The metric used to depend on exposure rather thannoise — the same pair of frames scaled down read 0.0163 % instead of 0.6172 %,so a night shot would have been told eight samples was plenty. The headline usedto quote the wrong rung's number for your setting. The ladder used to berendered with denoising off while your scene has it on. And the button thatapplies the result used to work on a report the add-on had already refused tostand behind.
Requires
Blender 4.2 or newer, and Cycles. No internet connection, no externallibraries; frames go to your system's temporary folder and nothing is left inyour project.
Whole deliveries, not one file at a time
- Check a folder. Point it at a delivery and every .blend inside is checked, each in its own background Blender. The scene you have open is not touched — not opened over, not saved, not changed. A file that hangs or crashes takes only its own process down; the other thirty-nine still get checked.
- A report you can send. One self-contained page: verdict first, one sentence a lead can read, then the numbers, then every finding with its address. No internet, no fonts to load, no dependencies — it opens in six months on a machine with no network, which is exactly when someone asks what you delivered.
- Runs without a person. A single command from your pipeline writes the report and returns a code: clean, something blocking, or could not be checked. Those are three different codes on purpose — a check that did not run has not passed, and a gate that treats them the same is worse than no gate at all.
- Remembers last time. Save today's numbers as a baseline and a later run tells you what moved, in which direction, and which findings are new. A measurement that went up is not automatically bad: each product says which way is worse for its own numbers.
- One verdict for the whole delivery. If you own more than one of these tools, one run puts every check you have over every file and answers the only question that matters before you send it: can this go out? Each file is opened once and all the checks see the same state of it.