Why German CSV exports break in Python, R, and Excel
During a university lab course, our group simulated tens of thousands of rows of data for an upcoming experiment and needed to turn that data into plots to explain it. The simulation wrote its output to a text file. The plotting code failed on the first run.
What was actually wrong
The data file used a comma as the decimal separator and a semicolon as the column separator, the standard convention on German Windows systems. A value like 23,5 meant twenty-three point five, not twenty-three and five. The plotting code, written for the international convention, expected a period for decimals and a comma between columns. It read 23,5;1013,2 as one garbled field instead of two clean numbers.
This is not a bug in either convention. Windows regional settings tie the decimal separator to the list separator: if the decimal separator is a comma, Excel switches the CSV delimiter to a semicolon so the two don't collide. Countries that use a comma for decimals, most of continental Europe, end up with semicolon-delimited CSVs. Countries that use a period for decimals, including the US and UK, end up with comma-delimited CSVs. Move a file between the two, and nothing lines up.
What we tried first
We tried pasting the data into AI chatbots and asking them to fix the format. That works for a short sample, but chatbot input fields have a context limit, and tens of thousands of rows of data plus the same number of rows of expected output does not fit. The file was too long before we even got to the fix.
We tried find-and-replace inside a text editor, swapping every comma for a period. This corrupts data as soon as any field legitimately contains a comma outside of a decimal number, and at that row count there was no realistic way to check by hand which replacements were safe.
Excel's Text Import Wizard (Data > From Text/CSV) can be told explicitly which character is the decimal separator and which is the delimiter, and it does work. But it requires knowing the wizard exists and which settings to change, and most people in the course didn't know to look for it.
Who ends up stuck
In the end, someone in the group who already knew how to script it wrote a short fix and passed it around. That's the actual pattern: people who can already write a script solve this for themselves in a few minutes and move on. Everyone else just asked around until someone handed them a working file. Nobody in that second group brought up Excel's import wizard, not because they'd tried it and didn't trust the result, but because they didn't know it existed.
What this tool does
The converter on this site parses the file structure first, decimal separators and column delimiters, before changing anything, so a comma inside a quoted text field (a sample name, for example) is never mistaken for a column break. It shows you what it detected and what it's about to do before you download anything. Everything runs in your browser; the file is never uploaded anywhere.
Warum deutsche CSV-Exporte in Python, R und Excel scheitern
Während eines Laborkurses an der Universität simulierte unsere Gruppe mehrere zehntausend Zeilen Daten für ein bevorstehendes Experiment und wollte diese Daten in Diagrammen darstellen. Die Simulation schrieb ihre Ausgabe in eine Textdatei. Das Plot-Skript scheiterte beim ersten Versuch.
Was tatsächlich falsch war
Die Datendatei verwendete ein Komma als Dezimaltrennzeichen und ein Semikolon als Spaltentrennzeichen, die Standardkonvention auf deutschen Windows-Systemen. Ein Wert wie 23,5 bedeutete dreiundzwanzig Komma fünf, nicht dreiundzwanzig und fünf. Das Plot-Skript, geschrieben für die internationale Konvention, erwartete einen Punkt für Dezimalzahlen und ein Komma zwischen den Spalten. Es las 23,5;1013,2 als ein einziges, verstümmeltes Feld statt zwei sauberer Zahlen.
Das ist in keiner der beiden Konventionen ein Fehler. Die Regionaleinstellungen von Windows koppeln das Dezimaltrennzeichen an das Listentrennzeichen: Ist das Dezimaltrennzeichen ein Komma, wechselt Excel das CSV-Trennzeichen zu einem Semikolon, damit sich beide nicht überschneiden. Länder, die ein Komma für Dezimalzahlen verwenden, der Großteil Kontinentaleuropas, erhalten dadurch semikolongetrennte CSV-Dateien. Länder, die einen Punkt für Dezimalzahlen verwenden, darunter die USA und Großbritannien, erhalten kommagetrennte CSV-Dateien. Verschiebt man eine Datei zwischen beiden Systemen, stimmt nichts mehr überein.
Was wir zuerst versucht haben
Wir haben versucht, die Daten in KI-Chatbots einzufügen und sie um eine Korrektur zu bitten. Das funktioniert bei einer kurzen Stichprobe, aber Eingabefelder von Chatbots haben ein Kontextlimit, und mehrere zehntausend Zeilen Eingabedaten plus die gleiche Anzahl an erwarteten Ausgabezeilen passen nicht hinein. Die Datei war zu lang, bevor wir überhaupt bei der Korrektur ankamen.
Wir haben Suchen-und-Ersetzen in einem Texteditor versucht und jedes Komma durch einen Punkt ersetzt. Das zerstört Daten, sobald ein Feld rechtmäßig ein Komma außerhalb einer Dezimalzahl enthält, und bei dieser Zeilenzahl gab es keine realistische Möglichkeit, von Hand zu prüfen, welche Ersetzungen sicher waren.
Der Textimport-Assistent von Excel (Daten > Aus Text/CSV) lässt sich explizit mitteilen, welches Zeichen das Dezimaltrennzeichen und welches das Trennzeichen ist, und das funktioniert tatsächlich. Aber dafür muss man wissen, dass der Assistent existiert und welche Einstellungen zu ändern sind, und die meisten im Kurs wussten nicht, danach zu suchen.
Wer am Ende hängen bleibt
Am Ende schrieb jemand aus der Gruppe, der bereits scripten konnte, eine kurze Korrektur und gab sie weiter. Das ist das eigentliche Muster: Wer schon programmieren kann, löst das Problem in ein paar Minuten selbst und macht weiter. Alle anderen fragten einfach herum, bis ihnen jemand eine funktionierende Datei gab. Niemand aus dieser zweiten Gruppe erwähnte den Textimport-Assistenten von Excel, nicht weil sie ihn ausprobiert und dem Ergebnis nicht vertraut hätten, sondern weil sie nicht wussten, dass es ihn gibt.
Was dieses Tool macht
Der Converter auf dieser Seite analysiert zuerst die Dateistruktur, Dezimaltrennzeichen und Spaltentrennzeichen, bevor er irgendetwas ändert, sodass ein Komma innerhalb eines Textfelds in Anführungszeichen (zum Beispiel ein Probenname) nie fälschlich als Spaltenumbruch interpretiert wird. Es zeigt Ihnen, was erkannt wurde und was als Nächstes passiert, bevor Sie irgendetwas herunterladen. Alles läuft in Ihrem Browser; die Datei wird nie hochgeladen.