
Working with JSON in Python: Serialization, Custom Encoders, and Pitfalls
JSON is the format Python programs use to talk to almost everything else: web APIs, JavaScript frontends, config files, message queues, and databases. Python's built-in json module handles the basics in two function calls, and for simple dicts of strings and numbers that's all you need.
The trouble starts at the edges. You try to serialize a datetime and get TypeError: Object of type datetime is not JSON serializable. A dict with integer keys comes back with string keys. A price stored as Decimal turns into a float and loses a cent. A 19-digit ID survives Python just fine and then gets silently rounded by the browser.
This post covers how Python types map to JSON, the formatting options worth knowing, how to serialize your own types with default and custom JSONEncoder classes, how to decode JSON back into rich types with hooks, error handling, and a list of pitfalls that come up in real projects. Examples run on Python 3.13. (Reading and writing JSON files, including JSON Lines, is covered in file handling in Python; this post focuses on the serialization itself.)
The Four Functions
| Function | Converts | Direction |
|---|---|---|
json.dumps(obj) | Python object to str | Serialize ("dump string") |
json.loads(s) | str (or bytes) to Python object | Deserialize ("load string") |
json.dump(obj, f) | Python object to a file | Serialize |
json.load(f) | File to Python object | Deserialize |
Everything in this post applies to both the string and file versions; they take the same keyword arguments.
import json
data = {"name": "Ada", "age": 36, "active": True, "score": 9.5,
"tags": ("admin", "dev"), "manager": None}
s = json.dumps(data)
print(s)
back = json.loads(s)
print(back)
print(back == data)
{"name": "Ada", "age": 36, "active": true, "score": 9.5, "tags": ["admin", "dev"], "manager": null}
{'name': 'Ada', 'age': 36, 'active': True, 'score': 9.5, 'tags': ['admin', 'dev'], 'manager': None}
False
The round trip isn't equal, and that's the first lesson: JSON has fewer types than Python, so some information is lost on the way out.
How Python Types Map to JSON
| Python | JSON | Comes back as |
|---|---|---|
dict | object | dict |
list, tuple | array | list |
str | string | str |
int, float | number | int or float |
True / False | true / false | bool |
None | null | None |
int- or str-based Enum (IntEnum, StrEnum) | number / string | plain int / str |
datetime, date, Decimal, set, bytes, UUID, dataclasses, other objects | not supported | TypeError |
Two things in that table cause most surprises.
Tuples become lists. That's why the round trip above wasn't equal: ("admin", "dev") came back as ["admin", "dev"]. If your code relies on tuples (as dict keys, or for hashing), convert them back after loading.
Object keys are always strings. JSON only allows string keys, so json.dumps converts int, float, bool, and None keys to strings, and they don't convert back:
d = {1: "one"}
restored = json.loads(json.dumps(d))
print(restored, restored == d) # {'1': 'one'} False
print(json.dumps({1: "a", 2.5: "b", None: "d"}))
# {"1": "a", "2.5": "b", "null": "d"}
Code that does restored[1] will raise KeyError. Other key types, like tuples, raise TypeError: keys must be str, int, float, bool or None, not tuple (pass skipkeys=True to drop them silently instead, though that's rarely what you want).
Formatting Options
Pretty Printing
print(json.dumps({"b": 1, "a": [1, 2]}, indent=2, sort_keys=True))
{
"a": [
1,
2
],
"b": 1
}
indent takes a number of spaces or a string like "\t". sort_keys=True gives you deterministic output, which matters for diffs, cache keys, snapshot tests, and signatures: two equal dicts built in different orders produce identical JSON.
Compact Output
By default, dumps puts a space after , and :. For the smallest payload, set separators:
print(json.dumps({"b": 1, "a": [1, 2]}, separators=(",", ":")))
# {"b":1,"a":[1,2]}
Non-ASCII Characters
By default, every non-ASCII character is escaped:
print(json.dumps({"city": "Zürich", "check": "✓"}))
print(json.dumps({"city": "Zürich", "check": "✓"}, ensure_ascii=False))
{"city": "Zürich", "check": "✓"}
{"city": "Zürich", "check": "✓"}
Both are valid JSON and decode to the same thing. The escaped form is safe to put through any channel that might mangle non-ASCII bytes; ensure_ascii=False is more readable and smaller for non-English text. If you use it while writing a file, make sure the file is opened with encoding="utf-8".
Serializing Custom Types with default
When json meets an object it doesn't know, it calls the function you pass as default, and serializes whatever that function returns. If the function can't handle the object either, it should raise TypeError.
Here's a default function that covers the types most applications run into:
import json
from datetime import UTC, date, datetime
from decimal import Decimal
from enum import Enum
from pathlib import Path
from uuid import UUID
def to_json(obj):
if isinstance(obj, (datetime, date)):
return obj.isoformat()
if isinstance(obj, Decimal):
return str(obj)
if isinstance(obj, (set, frozenset)):
return sorted(obj)
if isinstance(obj, (UUID, Path)):
return str(obj)
if isinstance(obj, Enum):
return obj.value
raise TypeError(f"Object of type {type(obj).__name__} is not JSON serializable")
class Color(Enum):
RED = 1
order = {
"id": UUID("12345678-1234-5678-1234-567812345678"),
"placed_at": datetime(2026, 9, 19, 15, 33, tzinfo=UTC),
"total": Decimal("19.99"),
"tags": {"gift", "express"},
"color": Color.RED,
}
print(json.dumps(order, default=to_json, indent=2))
{
"id": "12345678-1234-5678-1234-567812345678",
"placed_at": "2026-09-19T15:33:00+00:00",
"total": "19.99",
"tags": [
"express",
"gift"
],
"color": 1
}
A few design choices worth explaining:
- Datetimes become ISO 8601 strings. That's the format JavaScript's
Dateparses anddatetime.fromisoformat()reads back. Make sure they're timezone-aware first; a naive datetime produces a string with no offset, which the receiver will interpret however it likes. See dates and times with zoneinfo for why. - Decimals become strings, not floats. Converting
Decimal("19.99")to a float throws away the exactness that's the whole point of usingDecimal. A string preserves it. Many payment APIs use strings or integer cents for this reason. - Sets become sorted lists, so the output is deterministic.
- The final
raise TypeErroris important. IfdefaultreturnsNonefor unknown types, every unsupported object silently becomesnulland you'll only find out when data is missing downstream.
The Tempting Shortcut: default=str
You'll often see json.dumps(obj, default=str). It never fails, because everything has a str():
print(json.dumps({"dt": datetime(2026, 1, 1)}, default=str))
# {"dt": "2026-01-01 00:00:00"}
That's fine for logging and debugging. For anything another program will read, avoid it: str(datetime) uses a space instead of the ISO T separator (which some strict parsers reject), and a forgotten object becomes "<app.models.User object at 0x10a3f>" instead of an error. Be explicit about the types you support.
Custom JSONEncoder Subclasses
If you use the same serialization rules across a project, package them as a JSONEncoder subclass and override its default() method:
import json
from dataclasses import asdict, dataclass, field, is_dataclass
from datetime import UTC, datetime
from decimal import Decimal
class AppEncoder(json.JSONEncoder):
def default(self, obj):
if isinstance(obj, datetime):
return obj.isoformat()
if isinstance(obj, Decimal):
return str(obj)
if is_dataclass(obj) and not isinstance(obj, type):
return asdict(obj)
return super().default(obj) # raises TypeError for anything else
@dataclass
class LineItem:
sku: str
price: Decimal
qty: int = 1
@dataclass
class Order:
id: int
placed_at: datetime
items: list[LineItem] = field(default_factory=list)
o = Order(1, datetime(2026, 9, 19, 15, 33, tzinfo=UTC),
[LineItem("MUG-01", Decimal("12.50"), 2)])
print(json.dumps(o, cls=AppEncoder))
{"id": 1, "placed_at": "2026-09-19T15:33:00+00:00", "items": [{"sku": "MUG-01", "price": "12.50", "qty": 2}]}
Pass the class with cls=AppEncoder. Calling super().default(obj) at the end raises the standard TypeError for unsupported types.
dataclasses.asdict() recursively converts nested dataclasses, lists, and dicts into plain structures, and the encoder then handles the Decimal and datetime values that asdict leaves inside. The not isinstance(obj, type) check makes sure you don't try to serialize the dataclass class itself. For more on dataclasses, see dataclasses in Python.
default Isn't Called for Subclasses of Supported Types
default is only called for objects the encoder doesn't already know. If your type subclasses str, int, float, dict, or list, the encoder treats it as the base type and never consults your hook:
class Money(float):
pass
print(json.dumps({"m": Money(1.5)}, default=lambda o: "custom"))
# {"m": 1.5}
The same is why StrEnum and IntEnum members serialize as their values automatically while a plain Enum raises TypeError. If you need to customize how a built-in type is written (say, every float rounded to two places), you have to convert the data before calling dumps.
Decoding: Hooks for Richer Types
Going the other way, json.loads produces only dicts, lists, strings, numbers, booleans, and None. Three hooks let you customize that.
object_hook
object_hook is called with every decoded JSON object (as a dict), innermost first, and its return value replaces the dict. You can use it to turn known fields back into rich types:
import json
from datetime import datetime
def decode_dates(obj: dict) -> dict:
for key, value in obj.items():
if key.endswith("_at") and isinstance(value, str):
obj[key] = datetime.fromisoformat(value)
return obj
raw = '{"id": 1, "placed_at": "2026-09-19T15:33:00+00:00", "meta": {"shipped_at": "2026-09-20T09:00:00+00:00"}}'
print(json.loads(raw, object_hook=decode_dates))
{'id': 1, 'placed_at': datetime.datetime(2026, 9, 19, 15, 33, tzinfo=datetime.timezone.utc), 'meta': {'shipped_at': datetime.datetime(2026, 9, 20, 9, 0, tzinfo=datetime.timezone.utc)}}
Using a naming convention (*_at) is more reliable than trying to detect anything that "looks like a date", which will eventually misfire on a user-entered string.
parse_float for Exact Decimals
By default, JSON numbers with a decimal point become Python floats, with the usual binary floating-point rounding. parse_float lets you substitute Decimal:
import json
from decimal import Decimal
data = json.loads('{"price": 0.1, "qty": 3}', parse_float=Decimal)
print(data)
print(0.1 * 3, data["price"] * 3)
{'price': Decimal('0.1'), 'qty': 3}
0.30000000000000004 0.3
There's a matching parse_int hook, though you rarely need it, since Python ints are already arbitrary precision.
object_pairs_hook for Duplicate Keys
JSON technically allows duplicate keys, and Python keeps the last one without complaint:
print(json.loads('{"a": 1, "a": 2}')) # {'a': 2}
In a config file, that usually means a copy-paste mistake. object_pairs_hook receives the raw list of (key, value) pairs before they're turned into a dict, so you can reject duplicates:
import json
def no_duplicates(pairs: list[tuple[str, object]]) -> dict:
keys = [k for k, _ in pairs]
dupes = {k for k in keys if keys.count(k) > 1}
if dupes:
raise ValueError(f"duplicate keys: {sorted(dupes)}")
return dict(pairs)
json.loads('{"a": 1, "a": 2}', object_pairs_hook=no_duplicates)
# ValueError: duplicate keys: ['a']
If you need full validation of shapes and types, that's a job for a schema library like Pydantic rather than hand-written hooks; see data validation with Pydantic.
Handling Decode Errors
Invalid input raises json.JSONDecodeError, a subclass of ValueError, with attributes that pinpoint the problem:
import json
bad = '{"name": "Ada", "age": 36,}'
try:
json.loads(bad)
except json.JSONDecodeError as e:
print(e.msg, e.lineno, e.colno, e.pos)
Illegal trailing comma before end of object 1 26 25
(That specific trailing-comma message is new in Python 3.13; older versions say Expecting property name enclosed in double quotes.) Other common failures:
| Input | Error |
|---|---|
{'name': 'Ada'} | Expecting property name enclosed in double quotes (single quotes aren't JSON) |
"" (empty body) | Expecting value: line 1 column 1 (char 0) |
{"a": 1} {"b": 2} | Extra data (two documents; that's JSON Lines) |
The empty-string case is the one you'll see most from HTTP clients: an API returned an empty body or an HTML error page, and your code tried to parse it as JSON. Check the status code and content type before decoding, and catch JSONDecodeError at the boundary where untrusted input comes in.
json.loads also accepts bytes and bytearray directly (UTF-8, UTF-16, or UTF-32), so you don't need to .decode() a response body first. And the top-level value doesn't have to be an object: json.loads("42"), json.loads('"text"'), and json.loads("null") are all valid.
Pitfalls to Watch For
NaN and Infinity Aren't Valid JSON
Python happily writes them:
import json, math
print(json.dumps([math.nan, math.inf, -math.inf]))
# [NaN, Infinity, -Infinity]
That output is not valid JSON. Python and JavaScript's JSON.parse disagree here: Python reads it back fine, but JSON.parse("[NaN]") throws, and so do strict parsers in other languages. Pass allow_nan=False to fail fast instead:
json.dumps([math.nan], allow_nan=False)
# ValueError: Out of range float values are not JSON compliant: nan
Then decide explicitly how missing numbers should be represented, usually None.
Large Integers and JavaScript
Python serializes arbitrarily large ints exactly. JavaScript parses every JSON number as a 64-bit float, which can only represent integers exactly up to 2^53 - 1 (9,007,199,254,740,991). A database ID like 12345678901234567890 survives Python but is rounded in the browser, so the frontend ends up requesting the wrong record. This is why APIs like Twitter's historically sent IDs as both numbers and strings. If your IDs can exceed 2^53, send them as strings.
Floats Are Floats
json.dumps(1.1 + 2.2) gives 3.3000000000000003, because that's the actual float value. JSON doesn't fix floating-point arithmetic. Use Decimal with a string encoder and parse_float=Decimal on the way back in for money and other exact quantities.
Circular References
An object that contains itself can't be serialized:
a = []
a.append(a)
json.dumps(a)
# ValueError: Circular reference detected
This usually shows up with ORM objects or parent/child links (a Comment with a post that has a list of comments). Serialize a view of the data, such as IDs instead of back-references, rather than the object graph itself.
Embedding JSON in HTML
json.dumps doesn't escape <, >, or /, so a string containing </script> passes through unchanged. If you put JSON inside a <script> tag in a template, a malicious value can close the tag and inject HTML. Use your framework's safe helper (Django's json_script filter, for example) or replace < with < before embedding.
json.loads(json.dumps(x)) as a Deep Copy
It works for pure-JSON data, but it's slower than copy.deepcopy, and it silently converts tuples to lists and non-string keys to strings. Use copy.deepcopy() when you mean "copy".
Faster Alternatives
The standard library's json is implemented in C and fast enough for most applications. When serialization shows up in your profiler, third-party libraries such as orjson and msgspec are typically several times faster and natively support types like datetime, UUID, and dataclasses. Their APIs differ from the standard library's (for example, orjson.dumps() returns bytes), so check the docs before swapping them in.
For quick inspection from the command line, the standard library also ships a pretty-printer:
echo '{"b": 1, "a": [1, 2]}' | python -m json.tool --sort-keys
Conclusion
The json module handles the common case in one line, and almost every real-world problem lives at the type boundary. Remember what changes on the round trip (tuples to lists, keys to strings), give dumps a default function or a JSONEncoder subclass that converts your types explicitly and raises TypeError for anything else, and use object_hook and parse_float to bring rich types back on the way in.
Then guard against the pitfalls that cross language boundaries: NaN isn't JSON, big integers need to be strings for JavaScript, money shouldn't be a float, and JSONDecodeError will happen at every boundary where you don't control the input.


