Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Fix a Python CSV export that breaks on commas inside quoted fields

CI Python 3.12

A small order-export script built CSV rows by joining fields with str.join(",") and read them back with str.split(","). That works until a field itself contains a comma. Then every column after it shifts by one, silently, with no error. This repo shows the exact broken row, the two-line diagnosis, the fix using Python's standard csv module, and a test that fails on the old code and passes on the new one.

The bug

Fixture order 1001 has address: "12 Birkenweg, Berlin". Running the original script on it (before/export_orders.py) produces this CSV line:

1001,Anna Keller,12 Birkenweg, Berlin,49.99

and reading it back with the same naive split(",") gives:

{'order_id': '1001', 'customer': 'Anna Keller', 'address': '12 Birkenweg', 'total': ' Berlin'}

total is now the string " Berlin" instead of 49.99. The real total is gone. Rows with a comma later in the address, or a newline embedded in a field, corrupt even further, up to spilling into an extra line.

How it was found

  1. Ran the original script against a small JSON fixture: python before/export_orders.py fixtures/orders.json /tmp/out.csv.
  2. Opened the CSV and diffed it against the source JSON column by column.
  3. Found that only rows whose address contains a comma are affected, and confirmed the cause by reading export_orders.py: it joins and splits on a bare , with no quoting.

The fix

Replace the manual join/split with the stdlib csv module, which understands RFC 4180 quoting:

writer = csv.DictWriter(buf, fieldnames=FIELDS, quoting=csv.QUOTE_MINIMAL)
writer.writeheader()
writer.writerows(rows)

csv.DictWriter quotes any field containing a comma, quote, or newline, and csv.DictReader parses those quotes back out correctly, so no column ever shifts.

Run it in 30 seconds

git clone https://github.com/Mo-Ko/show-csv-parser-fix.git
cd show-csv-parser-fix
uv venv && uv pip install pytest   # or: python3 -m venv .venv && .venv/bin/pip install pytest
make test

# see the fix in action
.venv/bin/python -m src.cli export fixtures/orders.json out.csv
cat out.csv
.venv/bin/python -m src.cli import out.csv

Measured

$ pytest -q
...                                                                      [100%]
3 passed in 0.01s

The three tests: the before-script demonstrates the shifted-column bug on the fixture, rows_to_csv/csv_to_rows round-trip the fixture exactly (including the comma, quote, and newline fields), and an empty row list round-trips to a header-only file.

Port map

  • Core (src/orders_csv.py): rows_to_csv(rows) -> str and csv_to_rows(text) -> list[dict]. Pure functions, no file I/O, fully testable in memory.
  • Adapter (src/cli.py): reads JSON, calls the core, writes CSV, or reads CSV and prints JSON. All I/O lives here.
  • Adding a Google Sheets writer later is a new adapter that calls rows_to_csv/csv_to_rows; the core does not change.

What this deliberately does not do

  • No pandas, no third-party CSV libraries. The stdlib csv module is enough for this problem.
  • No schema validation, retries, or logging. This is the minimal correct fix, not a production pipeline.
  • Does not touch whatever downstream tool (spreadsheet macro, importer, etc.) consumes the CSV file; that is out of scope here.

FAQ

Why does Python's split(',') break on CSV files?

CSV is not just comma-separated text: RFC 4180 allows a field to contain a comma, a quote, or a newline as long as the whole field is wrapped in double quotes. A plain str.split(",") has no idea about quoting, so it splits inside quoted fields too, shifting every column after the split point.

Should I use pandas for this?

Not for a job this small. Pandas is a heavy dependency to pull in just to read and write a handful of columns, and it still needs the same quoting rules under the hood. The stdlib csv module handles RFC 4180 correctly with zero extra dependencies.

How do I handle quotes and newlines inside a CSV field?

Use csv.writer/csv.DictWriter with quoting=csv.QUOTE_MINIMAL (the default) to write, and csv.reader/csv.DictReader to read. They quote a field automatically whenever it contains a comma, a double quote, or a newline, and un-quote it correctly on the way back in, as shown in src/orders_csv.py.

Can this be scheduled or connected to Google Sheets?

Yes. The parsing logic in src/orders_csv.py has no file I/O, so it can be called from a scheduled job or from a new adapter that writes to the Google Sheets API instead of a local CSV file. The core functions would not need to change.

Mohsin Kokab, 2026. MIT.

About

Fix a Python CSV export that breaks on commas inside quoted fields: bug, diagnosis, fix, test. Showcase by Mohsin Kokab.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages