When the Stack Collapses: Diagnosing the Processing Mistakes That Quietly Destroy Deep-Sky Images
There is a particular frustration familiar to many astrophotographers: a night that felt productive — good seeing, solid guiding, three or four hours of clean exposures — that somehow yields a final image that looks worse than sessions half as long. The telescope performed. The mount tracked faithfully. The sky cooperated. And yet the stacked result is noisy, blotchy, or missing detail that should be there.
In the majority of these cases, the data is fine. The stack is not.
The integration stage — combining dozens or hundreds of individual frames into a single master image — is where astrophotography workflows are most vulnerable to compounding errors. Unlike a bad flat frame or an out-of-focus exposure, stacking mistakes are often invisible at the moment they occur and only reveal themselves in the finished image. Understanding how to diagnose them, and more importantly how to trace them back to specific decisions made during processing, is one of the most valuable skills a serious imager can develop.
The Rejection Algorithm Problem
Most stacking software — PixInsight, Siril, DeepSkyStacker, and their equivalents — offers multiple algorithms for identifying and excluding outlier pixels before combining frames. Sigma clipping, Winsorized sigma clipping, linear fit clipping, and various percentile-based methods each make different assumptions about the statistical distribution of your data.
The failure mode that appears most frequently in amateur workflows is applying an aggressive rejection threshold to a small number of frames. Sigma clipping at 2.0 sigma with 15 light frames, for instance, will flag and discard a significant fraction of legitimate signal as statistical outliers. The rejection algorithm cannot distinguish between a cosmic ray hit and a pixel that genuinely recorded more signal because the target briefly appeared brighter through a moment of excellent seeing. When too much data is rejected, the stack loses depth, noise increases, and fine structural detail — the kind that separates a good nebula image from a great one — simply disappears.
The diagnostic here is straightforward: generate a rejection map, which most modern stacking applications can produce alongside the master light. A healthy rejection map shows sparse, randomly distributed excluded pixels concentrated around satellite trails, cosmic ray events, and hot pixels. A rejection map that looks like diffuse noise across the entire frame, or that shows structured patterns, indicates that your threshold is too aggressive for your sample size.
As a general principle, sigma clipping becomes reliable only above approximately 20 to 25 frames. Below that count, Winsorized sigma clipping or linear fit clipping typically produces better-behaved results.
Weighting Schemes and Their Hidden Consequences
Frame weighting — assigning different levels of influence to individual exposures based on quality metrics — sounds like an obvious improvement over treating all frames equally. In practice, it introduces a new category of error if the weighting criterion is poorly chosen.
Background noise weighting, for example, rewards frames with lower measured sky background noise. This sounds sensible until you recognize that a frame acquired during a moment of excellent seeing, when the target's stars are tightly focused and the diffraction rings are sharp, may actually show slightly higher measured background noise due to increased sky glow from better atmospheric transparency. That frame — arguably your best data — gets penalized relative to frames from softer, hazier moments.
Signal-to-noise ratio weighting is generally more reliable for deep-sky targets, but even here the metric is computed differently across software platforms. Before trusting any automated weighting scheme, examine which frames are being upweighted and downweighted and verify that the ranking matches your subjective assessment of frame quality. If your sharpest, cleanest frames are consistently receiving lower weights, the algorithm is working against you.
Alignment Errors That Compound
Frame alignment is typically treated as a solved problem — the software finds the stars, registers the frames, and the process is complete. In reality, alignment quality degrades predictably under several conditions that are common in real observing sessions.
Large image rotations between frames, which occur when a German equatorial mount performs a meridian flip, challenge many alignment algorithms. If the software fails to account for the flip correctly, stars near the frame edges may be slightly misregistered while stars near the center align cleanly. The result in the final stack is a subtle but damaging gradient of sharpness — tight, well-resolved stars at the center giving way to faint elongation toward the corners.
A second alignment failure mode involves using too few reference stars, or reference stars that are too faint to centroid accurately. When the alignment reference set includes stars near the noise floor of individual frames, the computed transformation is noisy, and that noise accumulates across every frame in the stack. The symptom is a slight but measurable loss of resolution in the final image compared to individual exposures — a result that should be impossible if the stack is working correctly.
To diagnose alignment quality, blink through the registered frames before stacking and zoom in to the corners. If stars wander perceptibly between frames at the periphery, realign using only bright, well-exposed stars as references and repeat the comparison.
Separating Stack Failure from Data Failure
Before attributing a disappointing image to the stack, rule out the data itself. Examine five to ten individual calibrated frames directly. If they show strong, clean signal on your target with well-controlled background noise, the raw data is sound and the stack is the likely culprit. If individual frames look weak, grainy, or poorly calibrated — flat frames that don't match the vignetting pattern of the lights, or dark frames taken at a different temperature than the lights — the problem predates the stack entirely, and no amount of integration refinement will rescue it.
This distinction matters because the two failure modes require completely different remedies. A stack failure is correctable by revisiting your integration parameters. A data failure requires returning to the telescope.
The stack is not a safety net. It is a precision instrument, and it rewards the same careful attention as every other stage of the imaging process.