← back to the archiveCover illustration for “Reject duplicate JSON fields before validation”
TILday 130·today·Published ·by Andy Padia

Reject duplicate JSON fields before validation

In short: Reject repeated JSON fields while decoding in Python; a schema check on the resulting dictionary cannot recover values already overwritten.

python ▸ import json

def unique_object(pairs):
    result = {}
    for key, value in pairs:
        if key in result:
            raise ValueError("duplicate JSON field")
        result[key] = value
    return result

def load_unique(raw):
    return json.loads(raw, object_pairs_hook=unique_object)

assert load_unique('{"action":"read"}') == {"action": "read"}
bad = r'{"action":"read","\u0061ction":"write"}'
try:
    load_unique(bad)
except ValueError as error:
    assert str(error) == "duplicate JSON field"
    print("duplicate rejected")
else:
    raise AssertionError("duplicate accepted")

Two action fields entered Python's JSON decoder. One came out. In a synthetic Python 3.14.6 check on 8 October 2026, {"action":"read","action":"write"} became {"action":"write"} without an error. Setting strict=True did not change that result.

Reject duplicate JSON fields before validation turns into a dictionary-only exercise. My recommendation for agent-tool inputs is to catch repeated names during decoding, then validate the surviving structure and authorise the action separately. A clean-looking dictionary cannot tell you what the parser discarded.

Why duplicate JSON fields disappear

Python's JSON reference documents last-value behaviour for repeated names. Its strict option governs control characters inside strings; it is not a general promise to reject every problematic input.

RFC 8259, section 4 recommends unique names within each object and explains why duplicates hurt interoperability: receivers can retain the last value, reject the object or behave differently. This is not a newly discovered Python bug. It is a boundary decision your application should make explicitly.

Once the repeated name has collapsed into one dictionary entry, a validator receiving only that dictionary cannot inspect the missing pair. The point of intervention is earlier.

Reject duplicate JSON fields with a pairs hook

Python exposes object_pairs_hook for this job. The CPython 3.14.0 decoder source shows the hook receiving the collected pairs before the ordinary dictionary conversion. The check below rejects a repeated decoded name, even when both values match.

Save this as unique_json.py and run python3 unique_json.py. It uses only the standard library and synthetic strings:

import json

def unique_object(pairs):
    result = {}
    for key, value in pairs:
        if key in result:
            raise ValueError("duplicate JSON field")
        result[key] = value
    return result

def load_unique(raw):
    return json.loads(raw, object_pairs_hook=unique_object)

assert load_unique('{"action":"read"}') == {"action": "read"}
bad = r'{"action":"read","\u0061ction":"write"}'
try:
    load_unique(bad)
except ValueError as error:
    assert str(error) == "duplicate JSON field"
    print("duplicate rejected")
else:
    raise AssertionError("duplicate accepted")

The escaped spelling matters: \u0061ction decodes to action. Comparing the decoded names catches the collision without trying to scan raw JSON with a regular expression.

The local test set also covered duplicates inside a nested object and an object inside an array. All five duplicate fixtures were rejected; five distinct-key fixtures retained their expected values. Names repeated across separate objects remained valid. Those are functional checks, not throughput measurements or a test of a hosted model's output.

My judgement for an agent-tool input boundary

For a hypothetical document tool, I would put this check where the application first receives the raw argument text. Reject the ambiguous payload rather than guess which action the sender intended. Then run the normal type, range and permission checks. None of those jobs disappears because names are unique.

Check the framework boundary before adding the helper. If middleware has already parsed the body into a dictionary, applying this function afterwards is too late. Test the actual route with a duplicate-bearing request and verify that the tool is never invoked. That route-level test is a recommendation; the checks here ran locally, without a web framework or live tool.

Keep the rejection message generic. Echoing an entire malformed request into logs is a poor default when arguments might contain customer material. A reason code and a synthetic reproduction are often enough to start debugging.

What the duplicate check does not guarantee

This helper is deliberately narrow. It does not reject every non-standard numeric value Python accepts, impose payload-size limits or establish business correctness. It also cannot recover duplicates that an upstream service has already removed. Decide those other input policies separately instead of calling this a complete JSON security layer.

What's in it for you

  • Catch ambiguous tool arguments before dictionary conversion erases the evidence.
  • Add nested, escaped-name and identical-value duplicates to your input tests.

Validate what arrived before trusting the cleaner object your parser produced.

Sources

#python#testing#reliability
← older drop
AI research releases need a review budget

related drops

explore all 380 drops →
← back to the archiveday 130