Structured concurrency cancels siblings that already succeeded, not just the one that failed
In a structured-concurrency scope, one child throwing cancels every sibling immediately, and that cancellation is delivered to the others as an exception indistinguishable, to a generic catch block, from an ordinary failure of their own — so a sibling that had already produced real, salvageable output discards it as if its own work had gone wrong, when the only thing that actually failed was an unrelated task in the group.
Resurfaces when
- one task in a group of concurrent tasks fails and unrelated already-finished work also gets thrown away
- combining many downloaded or fetched items into one output as they complete
- wiring a producer and a stateful consumer as siblings in the same structured concurrency scope
- retrying an operation redoes far more work than the one piece that actually failed
- a catch block around a coroutine or task also catches its own cancellation and treats it like a normal failure
- designing a pipeline that downloads many items concurrently and joins them incrementally into one file
- a transient network blip on one item costs the progress of every other item
The lesson’s retrieval contract, weighted three times heavier than its body. A lesson that states when it applies doesn’t need a semantic search to find it.
The lesson
When one task in a structured-concurrency group throws, every sibling task is cancelled
immediately — not failed on its own terms, cancelled. This includes siblings that had already
done real, salvageable work that a later step was consuming incrementally (a writer that had
already flushed several successful items into a stateful, order-dependent output). A generic
catch (Exception) (or equivalent) wrapped around such a sibling will also catch the cancellation
signal, because in most runtimes "this scope is shutting down" is implemented as a subtype of the
ordinary exception hierarchy, not a separate channel. The sibling's cleanup path then runs exactly
as if it had failed by itself, discarding progress that had nothing to do with the actual failure.
Net effect: one transient failure in one task throws away correct, finished work from every other
task in the group, and whatever that group had built up has to be redone from scratch.
The failure that taught it
In a mobile app that downloads a series of remote files concurrently and incrementally muxes each into a single output container as it lands (order-dependent, one task consuming items produced by several others), the downloads and the muxer ran as sibling coroutines inside one structured concurrency block. A single download failing — one dropped connection among many, on an item unrelated to anything already muxed — cancelled the whole block, including the muxer mid-stream. The muxer's own error handling wrapped its loop in a general catch-and-clean-up, so it treated the cancellation the same as any other failure and discarded its output rather than distinguishing "my own work is fine, a sibling died." The underlying container format also could not be resumed once finalized or aborted, so every item that had already been correctly combined had to be redone from scratch on the next attempt — a full-series-length reset triggered by one item's transient, otherwise-harmless failure.
How to apply it
- Before wiring independent-but-order-dependent tasks into one structured-concurrency scope, ask specifically: "if task B throws, what does the cancellation of task A look like, distinct from A failing on its own?" A generic catch block almost never makes that distinction, and will treat cancellation-of-a-healthy-task the same as a real failure unless written to check for it.
- If a sibling task holds state or output that represents progress worth keeping (a partially written file, a partially filled cache, a partially sent batch), either give it a cancellation-aware path that preserves what is already done, or avoid putting it in the same cancellation scope as producers that can fail independently of it.
- Prefer fixing the root cause at the leaf — retry the actual flaky unit (a network call, a remote fetch) — so group-cancellation is rarely exercised in the first place, rather than trying to make every sibling gracefully survive being cancelled mid-work. Retrying one item is almost always simpler than making a stateful accumulator resumable.
- Recall this while designing a pipeline where one stage feeds items to another as they complete inside a single task group, or when a retry after a partial failure is doing far more work than the one piece that actually broke.
Where this claim comes from
- Lessonmedium confidence
Structured concurrency cancels siblings that already succeeded, not just the one that failed
- Distillationmechanism stated
Written from the mechanism, not the incident — which is what lets it transfer to code sharing nothing with the original.
- Scar1 occurrence
One recorded failure — weaker evidence, and ranked accordingly rather than presented as settled.
- Evidencenone recorded
No source records recorded — hand-written and migrated lessons predate the pipeline that captures them.
What happened when it was used
- 11
- retrieved
- 3
- acted on
- 0
- held up
- 0
- did not hold up
Applied 3 times. Outcomes move the ranking both ways, which is what makes this improve rather than just grow.