Type something to search...
Working with JSON in Python: Serialization, Custom Encoders, and Pitfalls

Working with JSON in Python: Serialization, Custom Encoders, and Pitfalls

JSON is the format Python programs use to talk to almost everything else: web APIs, JavaScript frontends, config files, message queues, and databases. Python's built-in json module handles the basics in two function calls, and for simple dicts of strings and numbers that's all you need.

The trouble starts at the edges. You try to serialize a datetime and get TypeError: Object of type datetime is not JSON serializable. A dict with integer keys comes back with string keys. A price stored as Decimal turns into a float and loses a cent. A 19-digit ID survives Python just fine and then gets silently rounded by the browser.

This post covers how Python types map to JSON, the formatting options worth knowing, how to serialize your own types with default and custom JSONEncoder classes, how to decode JSON back into rich types with hooks, error handling, and a list of pitfalls that come up in real projects. Examples run on Python 3.13. (Reading and writing JSON files, including JSON Lines, is covered in file handling in Python; this post focuses on the serialization itself.)

The Four Functions

FunctionConvertsDirection
json.dumps(obj)Python object to strSerialize ("dump string")
json.loads(s)str (or bytes) to Python objectDeserialize ("load string")
json.dump(obj, f)Python object to a fileSerialize
json.load(f)File to Python objectDeserialize

Everything in this post applies to both the string and file versions; they take the same keyword arguments.

import json

data = {"name": "Ada", "age": 36, "active": True, "score": 9.5,
        "tags": ("admin", "dev"), "manager": None}

s = json.dumps(data)
print(s)

back = json.loads(s)
print(back)
print(back == data)
{"name": "Ada", "age": 36, "active": true, "score": 9.5, "tags": ["admin", "dev"], "manager": null}
{'name': 'Ada', 'age': 36, 'active': True, 'score': 9.5, 'tags': ['admin', 'dev'], 'manager': None}
False

The round trip isn't equal, and that's the first lesson: JSON has fewer types than Python, so some information is lost on the way out.

How Python Types Map to JSON

PythonJSONComes back as
dictobjectdict
list, tuplearraylist
strstringstr
int, floatnumberint or float
True / Falsetrue / falsebool
NonenullNone
int- or str-based Enum (IntEnum, StrEnum)number / stringplain int / str
datetime, date, Decimal, set, bytes, UUID, dataclasses, other objectsnot supportedTypeError

Two things in that table cause most surprises.

Tuples become lists. That's why the round trip above wasn't equal: ("admin", "dev") came back as ["admin", "dev"]. If your code relies on tuples (as dict keys, or for hashing), convert them back after loading.

Object keys are always strings. JSON only allows string keys, so json.dumps converts int, float, bool, and None keys to strings, and they don't convert back:

d = {1: "one"}
restored = json.loads(json.dumps(d))
print(restored, restored == d)     # {'1': 'one'} False

print(json.dumps({1: "a", 2.5: "b", None: "d"}))
# {"1": "a", "2.5": "b", "null": "d"}

Code that does restored[1] will raise KeyError. Other key types, like tuples, raise TypeError: keys must be str, int, float, bool or None, not tuple (pass skipkeys=True to drop them silently instead, though that's rarely what you want).

Formatting Options

Pretty Printing

print(json.dumps({"b": 1, "a": [1, 2]}, indent=2, sort_keys=True))
{
  "a": [
    1,
    2
  ],
  "b": 1
}

indent takes a number of spaces or a string like "\t". sort_keys=True gives you deterministic output, which matters for diffs, cache keys, snapshot tests, and signatures: two equal dicts built in different orders produce identical JSON.

Compact Output

By default, dumps puts a space after , and :. For the smallest payload, set separators:

print(json.dumps({"b": 1, "a": [1, 2]}, separators=(",", ":")))
# {"b":1,"a":[1,2]}

Non-ASCII Characters

By default, every non-ASCII character is escaped:

print(json.dumps({"city": "Zürich", "check": "✓"}))
print(json.dumps({"city": "Zürich", "check": "✓"}, ensure_ascii=False))
{"city": "Zürich", "check": "✓"}
{"city": "Zürich", "check": "✓"}

Both are valid JSON and decode to the same thing. The escaped form is safe to put through any channel that might mangle non-ASCII bytes; ensure_ascii=False is more readable and smaller for non-English text. If you use it while writing a file, make sure the file is opened with encoding="utf-8".

Serializing Custom Types with default

When json meets an object it doesn't know, it calls the function you pass as default, and serializes whatever that function returns. If the function can't handle the object either, it should raise TypeError.

Here's a default function that covers the types most applications run into:

import json
from datetime import UTC, date, datetime
from decimal import Decimal
from enum import Enum
from pathlib import Path
from uuid import UUID


def to_json(obj):
    if isinstance(obj, (datetime, date)):
        return obj.isoformat()
    if isinstance(obj, Decimal):
        return str(obj)
    if isinstance(obj, (set, frozenset)):
        return sorted(obj)
    if isinstance(obj, (UUID, Path)):
        return str(obj)
    if isinstance(obj, Enum):
        return obj.value
    raise TypeError(f"Object of type {type(obj).__name__} is not JSON serializable")


class Color(Enum):
    RED = 1


order = {
    "id": UUID("12345678-1234-5678-1234-567812345678"),
    "placed_at": datetime(2026, 9, 19, 15, 33, tzinfo=UTC),
    "total": Decimal("19.99"),
    "tags": {"gift", "express"},
    "color": Color.RED,
}
print(json.dumps(order, default=to_json, indent=2))
{
  "id": "12345678-1234-5678-1234-567812345678",
  "placed_at": "2026-09-19T15:33:00+00:00",
  "total": "19.99",
  "tags": [
    "express",
    "gift"
  ],
  "color": 1
}

A few design choices worth explaining:

  • Datetimes become ISO 8601 strings. That's the format JavaScript's Date parses and datetime.fromisoformat() reads back. Make sure they're timezone-aware first; a naive datetime produces a string with no offset, which the receiver will interpret however it likes. See dates and times with zoneinfo for why.
  • Decimals become strings, not floats. Converting Decimal("19.99") to a float throws away the exactness that's the whole point of using Decimal. A string preserves it. Many payment APIs use strings or integer cents for this reason.
  • Sets become sorted lists, so the output is deterministic.
  • The final raise TypeError is important. If default returns None for unknown types, every unsupported object silently becomes null and you'll only find out when data is missing downstream.

The Tempting Shortcut: default=str

You'll often see json.dumps(obj, default=str). It never fails, because everything has a str():

print(json.dumps({"dt": datetime(2026, 1, 1)}, default=str))
# {"dt": "2026-01-01 00:00:00"}

That's fine for logging and debugging. For anything another program will read, avoid it: str(datetime) uses a space instead of the ISO T separator (which some strict parsers reject), and a forgotten object becomes "<app.models.User object at 0x10a3f>" instead of an error. Be explicit about the types you support.

Custom JSONEncoder Subclasses

If you use the same serialization rules across a project, package them as a JSONEncoder subclass and override its default() method:

import json
from dataclasses import asdict, dataclass, field, is_dataclass
from datetime import UTC, datetime
from decimal import Decimal


class AppEncoder(json.JSONEncoder):
    def default(self, obj):
        if isinstance(obj, datetime):
            return obj.isoformat()
        if isinstance(obj, Decimal):
            return str(obj)
        if is_dataclass(obj) and not isinstance(obj, type):
            return asdict(obj)
        return super().default(obj)   # raises TypeError for anything else


@dataclass
class LineItem:
    sku: str
    price: Decimal
    qty: int = 1


@dataclass
class Order:
    id: int
    placed_at: datetime
    items: list[LineItem] = field(default_factory=list)


o = Order(1, datetime(2026, 9, 19, 15, 33, tzinfo=UTC),
          [LineItem("MUG-01", Decimal("12.50"), 2)])
print(json.dumps(o, cls=AppEncoder))
{"id": 1, "placed_at": "2026-09-19T15:33:00+00:00", "items": [{"sku": "MUG-01", "price": "12.50", "qty": 2}]}

Pass the class with cls=AppEncoder. Calling super().default(obj) at the end raises the standard TypeError for unsupported types.

dataclasses.asdict() recursively converts nested dataclasses, lists, and dicts into plain structures, and the encoder then handles the Decimal and datetime values that asdict leaves inside. The not isinstance(obj, type) check makes sure you don't try to serialize the dataclass class itself. For more on dataclasses, see dataclasses in Python.

default Isn't Called for Subclasses of Supported Types

default is only called for objects the encoder doesn't already know. If your type subclasses str, int, float, dict, or list, the encoder treats it as the base type and never consults your hook:

class Money(float):
    pass


print(json.dumps({"m": Money(1.5)}, default=lambda o: "custom"))
# {"m": 1.5}

The same is why StrEnum and IntEnum members serialize as their values automatically while a plain Enum raises TypeError. If you need to customize how a built-in type is written (say, every float rounded to two places), you have to convert the data before calling dumps.

Decoding: Hooks for Richer Types

Going the other way, json.loads produces only dicts, lists, strings, numbers, booleans, and None. Three hooks let you customize that.

object_hook

object_hook is called with every decoded JSON object (as a dict), innermost first, and its return value replaces the dict. You can use it to turn known fields back into rich types:

import json
from datetime import datetime


def decode_dates(obj: dict) -> dict:
    for key, value in obj.items():
        if key.endswith("_at") and isinstance(value, str):
            obj[key] = datetime.fromisoformat(value)
    return obj


raw = '{"id": 1, "placed_at": "2026-09-19T15:33:00+00:00", "meta": {"shipped_at": "2026-09-20T09:00:00+00:00"}}'
print(json.loads(raw, object_hook=decode_dates))
{'id': 1, 'placed_at': datetime.datetime(2026, 9, 19, 15, 33, tzinfo=datetime.timezone.utc), 'meta': {'shipped_at': datetime.datetime(2026, 9, 20, 9, 0, tzinfo=datetime.timezone.utc)}}

Using a naming convention (*_at) is more reliable than trying to detect anything that "looks like a date", which will eventually misfire on a user-entered string.

parse_float for Exact Decimals

By default, JSON numbers with a decimal point become Python floats, with the usual binary floating-point rounding. parse_float lets you substitute Decimal:

import json
from decimal import Decimal

data = json.loads('{"price": 0.1, "qty": 3}', parse_float=Decimal)
print(data)
print(0.1 * 3, data["price"] * 3)
{'price': Decimal('0.1'), 'qty': 3}
0.30000000000000004 0.3

There's a matching parse_int hook, though you rarely need it, since Python ints are already arbitrary precision.

object_pairs_hook for Duplicate Keys

JSON technically allows duplicate keys, and Python keeps the last one without complaint:

print(json.loads('{"a": 1, "a": 2}'))   # {'a': 2}

In a config file, that usually means a copy-paste mistake. object_pairs_hook receives the raw list of (key, value) pairs before they're turned into a dict, so you can reject duplicates:

import json


def no_duplicates(pairs: list[tuple[str, object]]) -> dict:
    keys = [k for k, _ in pairs]
    dupes = {k for k in keys if keys.count(k) > 1}
    if dupes:
        raise ValueError(f"duplicate keys: {sorted(dupes)}")
    return dict(pairs)


json.loads('{"a": 1, "a": 2}', object_pairs_hook=no_duplicates)
# ValueError: duplicate keys: ['a']

If you need full validation of shapes and types, that's a job for a schema library like Pydantic rather than hand-written hooks; see data validation with Pydantic.

Handling Decode Errors

Invalid input raises json.JSONDecodeError, a subclass of ValueError, with attributes that pinpoint the problem:

import json

bad = '{"name": "Ada", "age": 36,}'
try:
    json.loads(bad)
except json.JSONDecodeError as e:
    print(e.msg, e.lineno, e.colno, e.pos)
Illegal trailing comma before end of object 1 26 25

(That specific trailing-comma message is new in Python 3.13; older versions say Expecting property name enclosed in double quotes.) Other common failures:

InputError
{'name': 'Ada'}Expecting property name enclosed in double quotes (single quotes aren't JSON)
"" (empty body)Expecting value: line 1 column 1 (char 0)
{"a": 1} {"b": 2}Extra data (two documents; that's JSON Lines)

The empty-string case is the one you'll see most from HTTP clients: an API returned an empty body or an HTML error page, and your code tried to parse it as JSON. Check the status code and content type before decoding, and catch JSONDecodeError at the boundary where untrusted input comes in.

json.loads also accepts bytes and bytearray directly (UTF-8, UTF-16, or UTF-32), so you don't need to .decode() a response body first. And the top-level value doesn't have to be an object: json.loads("42"), json.loads('"text"'), and json.loads("null") are all valid.

Pitfalls to Watch For

NaN and Infinity Aren't Valid JSON

Python happily writes them:

import json, math

print(json.dumps([math.nan, math.inf, -math.inf]))
# [NaN, Infinity, -Infinity]

That output is not valid JSON. Python and JavaScript's JSON.parse disagree here: Python reads it back fine, but JSON.parse("[NaN]") throws, and so do strict parsers in other languages. Pass allow_nan=False to fail fast instead:

json.dumps([math.nan], allow_nan=False)
# ValueError: Out of range float values are not JSON compliant: nan

Then decide explicitly how missing numbers should be represented, usually None.

Large Integers and JavaScript

Python serializes arbitrarily large ints exactly. JavaScript parses every JSON number as a 64-bit float, which can only represent integers exactly up to 2^53 - 1 (9,007,199,254,740,991). A database ID like 12345678901234567890 survives Python but is rounded in the browser, so the frontend ends up requesting the wrong record. This is why APIs like Twitter's historically sent IDs as both numbers and strings. If your IDs can exceed 2^53, send them as strings.

Floats Are Floats

json.dumps(1.1 + 2.2) gives 3.3000000000000003, because that's the actual float value. JSON doesn't fix floating-point arithmetic. Use Decimal with a string encoder and parse_float=Decimal on the way back in for money and other exact quantities.

Circular References

An object that contains itself can't be serialized:

a = []
a.append(a)
json.dumps(a)
# ValueError: Circular reference detected

This usually shows up with ORM objects or parent/child links (a Comment with a post that has a list of comments). Serialize a view of the data, such as IDs instead of back-references, rather than the object graph itself.

Embedding JSON in HTML

json.dumps doesn't escape <, >, or /, so a string containing </script> passes through unchanged. If you put JSON inside a <script> tag in a template, a malicious value can close the tag and inject HTML. Use your framework's safe helper (Django's json_script filter, for example) or replace < with < before embedding.

json.loads(json.dumps(x)) as a Deep Copy

It works for pure-JSON data, but it's slower than copy.deepcopy, and it silently converts tuples to lists and non-string keys to strings. Use copy.deepcopy() when you mean "copy".

Faster Alternatives

The standard library's json is implemented in C and fast enough for most applications. When serialization shows up in your profiler, third-party libraries such as orjson and msgspec are typically several times faster and natively support types like datetime, UUID, and dataclasses. Their APIs differ from the standard library's (for example, orjson.dumps() returns bytes), so check the docs before swapping them in.

For quick inspection from the command line, the standard library also ships a pretty-printer:

echo '{"b": 1, "a": [1, 2]}' | python -m json.tool --sort-keys

Conclusion

The json module handles the common case in one line, and almost every real-world problem lives at the type boundary. Remember what changes on the round trip (tuples to lists, keys to strings), give dumps a default function or a JSONEncoder subclass that converts your types explicitly and raises TypeError for anything else, and use object_hook and parse_float to bring rich types back on the way in.

Then guard against the pitfalls that cross language boundaries: NaN isn't JSON, big integers need to be strings for JavaScript, money shouldn't be a float, and JSONDecodeError will happen at every boundary where you don't control the input.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading