
Buffer strict zip before writing results
In short: Buffer a finite strict zip before writing results; length checking happens during iteration and cannot undo earlier sink writes.
python ▸ def collect_pairs(ids, values):
# Only for finite batches that comfortably fit in memory.
return list(zip(ids, values, strict=True))
def write_after_check(ids, values, output):
pairs = collect_pairs(ids, values)
for pair in pairs:
output.append(pair) # Synthetic sink, not a database.
written = []
try:
write_after_check(["a", "b", "c"], [10, 20], written)
except ValueError:
assert written == []
else:
raise AssertionError("unequal inputs accepted")
write_after_check(["a", "b"], [10, 20], written)
assert written == [("a", 10), ("b", 20)]
print("mismatch: zero sink writes; equal inputs: two")Three IDs, two values, and two writes before the error. That was the result of a synthetic Python 3.14.6 check on 10 October 2026, even with zip(strict=True). The length check worked. It simply arrived after the loop had already used the matching pairs.
Buffer strict zip before writing results when the whole finite batch must have matching lengths. My rule for small batches is to separate collecting valid pairs from touching the destination. A stricter iterator helps detect bad input; it does not make the surrounding operation atomic.
Why strict zip can fail after a write
Python's built-in function reference describes zip as lazy: it produces pairs when consumed. Ordinary zip stops at the shortest input. With strict=True, it raises ValueError when one input runs out before another.
PEP 618, which introduced the option for Python 3.10, locates the error at the point where iteration would otherwise stop. That timing matters. A loop can receive several valid pairs and append, upload or insert each one before asking for the pair that exposes the mismatch.
Constructing the iterator is therefore not a validation step. Neither is consuming one pair and breaking. The unseen tail is still unseen. Catching the eventual exception prevents an unhandled failure, but it does not reverse work the loop already performed.
Buffer strict zip before the sink loop
For a finite batch that comfortably fits in memory, collect all pairs first. Save this as zip_check.py and run python3 zip_check.py on Python 3.10 or later. The destination below is deliberately just a list:
def collect_pairs(ids, values):
# Only for finite batches that comfortably fit in memory.
return list(zip(ids, values, strict=True))
def write_after_check(ids, values, output):
pairs = collect_pairs(ids, values)
for pair in pairs:
output.append(pair) # Synthetic sink, not a database.
written = []
try:
write_after_check(["a", "b", "c"], [10, 20], written)
except ValueError:
assert written == []
else:
raise AssertionError("unequal inputs accepted")
write_after_check(["a", "b"], [10, 20], written)
assert written == [("a", 10), ("b", 20)]
print("mismatch: zero sink writes; equal inputs: two")
The important move is the assignment before the sink loop. If materialisation fails, execution never reaches that loop. Once collection succeeds, the two input streams have been exhausted together for that batch.
The synthetic Python batch check
The local checks exercised both mismatch directions, equal inputs and an empty batch. These selected results show the boundary:
| Strategy and input lengths | Sink writes | Outcome |
|---|---|---|
| Direct strict loop, 3 and 2 | 2 | ValueError |
| Buffered strict zip, 3 and 2 | 0 | ValueError |
| Buffered strict zip, 2 and 2 | 2 | Completed |
| Buffered strict zip, 0 and 0 | 0 | Completed |
Synthetic Python 3.14.6 results using an in-memory list; no database or network calls.
These are behavioural checks, not a performance benchmark. The separate early-stop check returned its first pair without checking the unequal tail. Add that case if your production loop contains a break or returns early.
My judgement for an embedding batch
For a hypothetical embedding pipeline, I would validate the returned IDs and vectors before starting destination writes. Buffering catches unequal counts, but equal counts can still contain reordered or incorrect pairs. Validate identity and vector shape separately rather than treating successful zip exhaustion as proof of correspondence.
There is another limit: buffering only delays the sink loop. Input generators may perform their own side effects while being consumed. A later destination write may also fail after earlier writes succeeded. This example does not provide rollback, and no real vector-store integration was tested.
For large or unbounded streams, collecting everything is the wrong design. Choose an explicit batch boundary and a destination-specific transaction or staging strategy. Decide whether the guarantee applies to one batch or the entire job; don't quietly promise the latter after checking only the former.
What's in it for you
- Catch unequal finite inputs before their matching prefix reaches the destination.
- Test both the raised error and the number of writes already performed.
Check the complete batch before writing its first pair—not after the iterator discovers its last mismatch.
Sources
- Python — Built-in functions, zip, Python 3.14, accessed 10 October 2026.
- Python — PEP 618, Optional length-checking for zip, created 1 May 2020; Python 3.10.


