
File Handling in Python: Reading and Writing Text, CSV, and JSON
Reading and writing files is one of the first things most people do in Python, and one of the places where small mistakes quietly cause trouble later: a script that works on your Mac but garbles accented characters on a Windows server, a CSV export with blank lines between every row, a config file left half-written after a crash.
The fixes are all small once you know them. This post walks through the built-in open() function and its modes, encodings, reading strategies for small and large files, the csv module for spreadsheets and exports, reading and writing JSON files, and a pattern for writing files safely so a crash can't corrupt them.
Everything here uses only the standard library and runs on Python 3.13.
Opening Files with open() and with
Every file operation starts with open(), which returns a file object. Always use it inside a with block:
with open("notes.txt", "w", encoding="utf-8") as f:
f.write("First line\n")
f.write("Second line\n")
The with statement is a context manager: it guarantees the file is closed when the block ends, even if an exception is raised inside it. Closing matters because writes are buffered in memory and only guaranteed to reach the disk when the file is flushed or closed. Without with, you have to remember f.close() on every code path, including error paths.
Once a file is closed, you can't use it anymore:
f.read()
# ValueError: I/O operation on closed file.
File Modes
The second argument to open() is the mode:
| Mode | Meaning | If the file exists | If it doesn't |
|---|---|---|---|
"r" | Read (default) | Opens it | FileNotFoundError |
"w" | Write | Truncates it to empty | Creates it |
"a" | Append | Writes go to the end | Creates it |
"x" | Exclusive create | FileExistsError | Creates it |
"r+" | Read and write | Opens without truncating | FileNotFoundError |
Add "b" for binary mode ("rb", "wb"), where you read and write bytes instead of str. Text mode ("t") is the default.
The one to be careful with is "w": opening an existing file in write mode erases it immediately, before you've written anything. If you mean "create a new file, but never clobber an existing one", use "x":
try:
with open("notes.txt", "x", encoding="utf-8") as f:
f.write("draft")
except FileExistsError as e:
print("FileExistsError:", e)
FileExistsError: [Errno 17] File exists: 'notes.txt'
Always Pass an Encoding
In text mode, Python converts between bytes on disk and str in memory using an encoding. If you don't pass one, it uses a platform-dependent default. On macOS and Linux that's almost always UTF-8. On Windows it has historically been a legacy code page like cp1252, so the same script can write a file that reads back fine on one machine and raises UnicodeDecodeError (or produces garbage) on another.
The fix is to always write encoding="utf-8" explicitly. Python 3.10 added an opt-in EncodingWarning (run with -X warn_default_encoding) that flags every open() call missing an encoding, which is a quick way to audit a codebase. PEP 686 makes UTF-8 the default in Python 3.15, but being explicit costs nothing and keeps your code correct on every version.
When a file isn't in the encoding you expect, decoding fails:
with open("latin.txt", encoding="utf-8") as f:
f.read()
# UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 3: unexpected end of data
You have three options:
- Use the right encoding if you know it:
open("latin.txt", encoding="latin-1"). - Tolerate bad bytes with
errors="replace"(substitutes�) orerrors="ignore"(drops them). Good for logs, bad for data you'll write back. - Open in binary mode and handle the bytes yourself.
One special case: files saved by Excel and some Windows tools start with a byte order mark (BOM). Read them with encoding="utf-8-sig", which strips the BOM if present. Otherwise your first column header becomes "id" instead of "id", and row["id"] raises KeyError. For the full story on bytes and encodings, see working with bytes, encoding, and Unicode.
Reading Text Files
There are several ways to read, and the right one depends on the file's size and what you're doing with it.
Read Everything at Once
with open("notes.txt", encoding="utf-8") as f:
content = f.read() # one string
with open("notes.txt", encoding="utf-8") as f:
lines = f.read().splitlines() # list of lines without "\n"
Simple and fine for config files, templates, and anything that comfortably fits in memory. splitlines() is usually nicer than readlines(), because readlines() keeps the trailing "\n" on each line.
Iterate Line by Line
A file object is an iterator over its lines. This is the idiomatic way to process a file, and it reads one line at a time no matter how big the file is:
with open("notes.txt", encoding="utf-8") as f:
for line_no, line in enumerate(f, start=1):
print(line_no, line.rstrip("\n"))
1 First line
2 Second line
Each line includes its trailing newline, so strip it with rstrip("\n") (or strip() if you also want to drop surrounding whitespace). This pattern works on multi-gigabyte log files with flat memory use.
readline and readlines
f.readline() returns the next line (or "" at the end of the file), which is useful for consuming a header before looping over the rest. f.readlines() returns all remaining lines as a list.
with open("data.txt", encoding="utf-8") as f:
header = f.readline().strip()
for line in f: # continues after the header
...
Reading in Chunks
For binary files, or text without useful line breaks, read fixed-size chunks. The walrus operator keeps the loop compact:
import hashlib
digest = hashlib.sha256()
with open("backup.tar.gz", "rb") as f:
while chunk := f.read(64 * 1024):
digest.update(chunk)
print(digest.hexdigest())
f.read(n) returns at most n characters (or bytes in binary mode) and an empty value at the end of the file, which ends the loop.
Writing Text Files
write() takes a single string and doesn't add a newline for you. writelines() takes an iterable of strings and also doesn't add newlines, despite the name. You can also point print() at a file:
with open("notes.txt", "w", encoding="utf-8") as f:
f.write("First line\n")
print("Second line", file=f) # print adds the newline
f.writelines(["Third\n", "Fourth\n"])
To add to the end of an existing file, such as a simple log, open it in "a" mode:
with open("notes.txt", "a", encoding="utf-8") as f:
f.write("Appended\n")
Newline Handling
In text mode, Python translates line endings for you. When reading, \r\n (Windows) and \r are converted to \n. When writing, \n is converted to the platform's native ending (\r\n on Windows). That's usually what you want, with one big exception: the csv module, covered next.
CSV Files with the csv Module
CSV looks simple enough to parse with line.split(","), until a field contains a comma, a quote, or a line break. The csv module handles all the quoting rules for you.
The newline="" Rule
Always open CSV files with newline="", for both reading and writing:
open("products.csv", "w", newline="", encoding="utf-8")
The csv writer emits \r\n line endings itself (that's what the CSV spec says). Without newline="", on Windows Python then translates the \n in that into \r\n again, producing \r\r\n, which spreadsheet apps show as a blank row between every record. When reading, newline="" lets the csv module correctly handle quoted fields that contain line breaks. It's harmless on macOS and Linux, so write it everywhere.
Writing Rows
import csv
rows = [
["sku", "name", "price"],
["MUG-01", "Coffee mug", "12.50"],
["LAMP-04", 'Desk lamp, "LED"', "65.00"],
]
with open("products.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerows(rows)
sku,name,price
MUG-01,Coffee mug,12.50
LAMP-04,"Desk lamp, ""LED""",65.00
The writer quoted the field that contained a comma and doubled the quotes inside it, exactly as the format requires. Use writerow() for one row at a time.
Reading Rows
with open("products.csv", newline="", encoding="utf-8") as f:
reader = csv.reader(f)
header = next(reader)
for row in reader:
print(row)
['MUG-01', 'Coffee mug', '12.50']
['LAMP-04', 'Desk lamp, "LED"', '65.00']
Every value comes back as a string. CSV has no types, so converting "12.50" to a number is your job.
DictReader and DictWriter
Index-based access (row[2]) breaks as soon as someone reorders the columns. DictReader uses the header row as keys, so each row becomes a dict:
with open("products.csv", newline="", encoding="utf-8") as f:
for record in csv.DictReader(f):
print(record["name"], float(record["price"]))
Coffee mug 12.5
Desk lamp, "LED" 65.0
Because the reader is lazy, aggregations stay memory-efficient:
with open("products.csv", newline="", encoding="utf-8") as f:
total = sum(float(r["price"]) for r in csv.DictReader(f))
DictWriter is the mirror image. You give it the column order up front and pass dicts:
import csv
people = [
{"name": "Ada", "email": "ada@example.com", "age": 36},
{"name": "Grace", "email": "grace@example.com", "age": 45},
]
with open("people.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["name", "email", "age"])
writer.writeheader()
writer.writerows(people)
name,email,age
Ada,ada@example.com,36
Grace,grace@example.com,45
Non-string values like 36 are converted with str(). If a dict has a key that isn't in fieldnames, DictWriter raises ValueError: dict contains fields not in fieldnames: 'extra'. Pass extrasaction="ignore" to drop unknown keys instead. Missing keys are written as empty strings (configurable with restval).
Delimiters, Dialects, and Quoting
Not every "CSV" uses commas. European exports often use semicolons, and TSV files use tabs:
with open("export.csv", newline="", encoding="utf-8") as f:
rows = list(csv.reader(f, delimiter=";"))
with open("data.tsv", newline="", encoding="utf-8") as f:
rows = list(csv.reader(f, dialect="excel-tab"))
If you don't know the format in advance, csv.Sniffer().sniff(sample) can guess the dialect from a sample of the file. It's a heuristic, so use it for one-off imports rather than production pipelines where you can agree on a format.
The quoting parameter controls when fields get quotes. csv.QUOTE_MINIMAL (the default) quotes only when needed. csv.QUOTE_NONNUMERIC quotes every non-numeric field on write and, on read, converts unquoted fields to float:
with open("q.csv", "w", newline="", encoding="utf-8") as f:
csv.writer(f, quoting=csv.QUOTE_NONNUMERIC).writerow(["Ada", 36, 1.5])
# file contains: "Ada",36,1.5
with open("q.csv", newline="", encoding="utf-8") as f:
print(next(csv.reader(f, quoting=csv.QUOTE_NONNUMERIC)))
# ['Ada', 36.0, 1.5]
For heavy analysis of tabular data, a DataFrame library is the better tool; see how to use Python for data analysis. The csv module is ideal for streaming, simple imports and exports, and scripts with no dependencies.
JSON Files
The json module reads and writes JSON files with two functions: json.dump() writes an object to a file, and json.load() reads one back. (The versions ending in s, dumps() and loads(), work with strings instead of files.)
import json
config = {"debug": True, "port": 8000, "hosts": ["a.example", "b.example"], "name": "Café"}
with open("config.json", "w", encoding="utf-8") as f:
json.dump(config, f, indent=2, ensure_ascii=False)
with open("config.json", encoding="utf-8") as f:
loaded = json.load(f)
print(loaded == config, loaded["hosts"])
True ['a.example', 'b.example']
The written file looks like this:
{
"debug": true,
"port": 8000,
"hosts": ["a.example", "b.example"],
"name": "Café"
}
indent=2 makes the file human-readable. ensure_ascii=False writes é as-is instead of the escape é; that's why the explicit encoding="utf-8" matters here too.
json.load() reads the whole file into memory at once, which is fine for configs and API fixtures. Types that JSON doesn't support (datetimes, sets, Decimal, your own classes) need a custom encoder; that's covered in detail in working with JSON in Python.
JSON Lines for Large or Append-Only Data
For logs, events, and datasets too big to load at once, use JSON Lines (.jsonl): one complete JSON object per line. You can append to it and stream through it line by line, combining the techniques from earlier in this post:
import json
events = [{"type": "click", "x": 1}, {"type": "view", "page": "/"}]
with open("events.jsonl", "a", encoding="utf-8") as f:
for event in events:
f.write(json.dumps(event) + "\n")
with open("events.jsonl", encoding="utf-8") as f:
for line in f:
if line.strip():
print(json.loads(line)["type"])
Writing Files Safely
If your program crashes, or the machine loses power, halfway through writing config.json, you're left with a truncated, invalid file, and the previous good version is gone because "w" erased it. For files that matter, write to a temporary file in the same directory, then swap it into place with os.replace():
# app/fileutils.py
import os
import tempfile
from pathlib import Path
def write_atomic(path: str | Path, text: str) -> None:
path = Path(path)
fd, tmp = tempfile.mkstemp(dir=path.parent, prefix=path.name, suffix=".tmp")
try:
with os.fdopen(fd, "w", encoding="utf-8") as f:
f.write(text)
f.flush()
os.fsync(f.fileno())
os.replace(tmp, path)
except BaseException:
os.unlink(tmp)
raise
os.replace() is atomic on the same filesystem on both POSIX and Windows: other readers see either the complete old file or the complete new one, never a partial write. os.fsync() forces the data to disk before the rename. The temporary file has to be in the same directory (not /tmp) so it's on the same filesystem.
Common Errors and What They Mean
| Error | Usual cause |
|---|---|
FileNotFoundError | Wrong path, or a relative path resolved against an unexpected working directory |
PermissionError | No permission, or the path is a directory, or (on Windows) another program has the file locked |
IsADirectoryError | You passed a directory to open() |
UnicodeDecodeError | The file isn't in the encoding you specified |
ValueError: I/O operation on closed file | Using f after its with block ended |
The first one deserves a note: relative paths are resolved against the current working directory, not the script's location. Run the script from a different folder and open("data.csv") looks in the wrong place. Building paths relative to Path(__file__).parent fixes that; the pathlib guide covers path handling in depth.
Conclusion
Most file-handling bugs come down to a handful of habits. Use with so files always close. Pass encoding="utf-8" every time, and utf-8-sig for files from Excel. Iterate over the file object instead of reading huge files into memory. Open CSV files with newline="" and use DictReader/DictWriter so columns are named. Use json.dump()/json.load() for structured data and JSON Lines when it needs to stream. And when a half-written file would be a problem, write to a temp file and os.replace() it into place.


