This article was originally written in Japanese and translated into English with AI assistance. Please note that some expressions may carry nuances from the original Japanese.
Series: utf8conv Development Journal (Part 4)
At the end of last time, utf8conv had become “the smallest thing that’s usable for now (an MVP).” However, the candidate character encodings it tried when converting to UTF-8 were only five: the four Japanese-related ones (Shift-JIS, CP932, EUC-JP, ISO-2022-JP) and the final catch-all, latin-1 (UTF-8 is outside the candidates. Before trying the candidates, it first checks separately whether the file “can be read as UTF-8”).
This time continues the same March 8 — the fourth round of work on that day. I did four things. I widened the candidates to 16 kinds. I exported utf8conv in the form of a Mac app (.app). I rewrote the requirements document as Rev.2. And I generated 518 test files all at once.
Let me state up front the one thing I want to say in this installment. On the day I added candidates, it became clear that “even for the same file, the answer changes depending on the order in which you try.” I did not fix that on the detection side; I wrote it into what the tests demand, and into the “Limitations” section of the requirements document.
Adding 11 to Make 16
If utf8conv’s future expansion overseas is to be kept in view, I want it to handle not only Japanese but also Chinese, Korean, and Eastern and Western European character encodings. With that in mind, I decided to do “international character encoding support.” I added 11 kinds. For Chinese, GBK, GB2312, and Big5; for Korean, EUC-KR and CP949; for Cyrillic (Russian and others), CP1251 and KOI8-R; for Eastern European languages, CP1250 and ISO-8859-2; for Western European languages, CP1252 and ISO-8859-1. Together with the original five, the candidates to try came to 16.
The ordering is: the Japanese-related ones first, followed by Chinese, Korean, Cyrillic, Eastern European, and Western European, with latin-1 placed last. utf8conv tries them in this order from the top and takes the first one that can read the file as the answer.
The Same Byte Sequence Can Be Read Two Ways
When I added the candidates, the tests caught on two things.
The first was that text written in Chinese (GBK) was detected as EUC-JP (Japanese).
Let me explain a little. The contents of a file are, when you get right down to it, a sequence of numbers (a byte sequence). For example, if you save the Chinese word “中文” in GBK, the contents look like this.
d6 d0 ce c4
This is a way of writing called hexadecimal, which expresses numbers using 16 symbols: 0 through 9 plus a through f. One byte on a computer can be written in exactly these two characters, so it has become the custom for showing byte sequences. Here, just read it as “four numbers in a row.”
If you read this byte sequence as EUC-JP, no error occurs. It reads as the two kanji characters “嶄猟.”
| Chinese (saved in GBK) | Byte sequence | Read as EUC-JP |
|---|---|---|
| 中文 | d6 d0 ce c4 | 嶄猟 |
| 日本 | c8 d5 b1 be | 晩云 |
| 北京 | b1 b1 be a9 | 臼奨 |
GBK and EUC-JP overlap in the range of bytes they use. Since neither produces an error, if you judge only by “whether it could be read,” whichever is tried first wins. utf8conv tries EUC-JP before GBK, so the Chinese file was answered with “this is Japanese.”
What I fixed at this point was not the detection, but the test. I changed the result the Chinese test demands from “detected as GBK” to “detected as something that is not UTF-8, and the contents can be read out.”
The second was the test with the four bytes 80 81 82 83. Until then it had been read as latin-1, but now it came to be read first by the newly added CP1251 (Cyrillic). Here too, I changed the test, to “it is fine as long as one of the final catch-alls can read it.”
In that day’s work, I sorted these two out as follows. In detection where the order of trying decides the answer, every time you add an encoding, it affects the tests that came before. As long as we do not use an additional part that detects statistically, checking “that it can be read out correctly as something that is not UTF-8” holds up better than “hitting the exact encoding name.”
Looking back now, the second one had a sequel. Three of the added ones — KOI8-R, ISO-8859-2, and ISO-8859-1 — can, like latin-1, read any byte sequence whatsoever. One of them answers “I could read it” first, so latin-1, at the very end, never gets its turn anymore. At this point, the catch-all was not working as a catch-all. How that story was settled, I will cover in a future blog post.
Exporting as a Mac App
The next task was exporting utf8conv in the form of a Mac .app, using a tool called PyInstaller. A .app is that thing with an icon that sits in the Mac’s Applications folder. Once this exists, just double-clicking in the Finder opens the utf8conv screen.
This work of assembling the code you have written into a single bundle you can hand to someone is called a build. The finished bundle, for its part, is a distributable. These two words will come up again and again from here on.
PyInstaller went into the category of tools used only during development, the same as pytest and the like.
The Tcl/Tk story also had a sequel. As I wrote in Part 1 of this series, on day one the Python that uv provides did not include Tcl/Tk (the part tkinter uses to draw the screen), so I switched to the Homebrew version of Python to get it running. PyInstaller found Tcl/Tk on its own from that Homebrew version of Python and packed it into the .app along with everything else.
On this day, dist/UTF-8 Converter.app came into being. I added the export command to the README.
Rewriting the Requirements Document as Rev.2
An hour later, I rewrote the requirements document from Rev.1 to Rev.2. Compared with Rev.1, written the day before, there were four main changes.
The first is the background. To the list of situations where mojibake occurs, I added “cases that are not resolved even when you tell an editor such as VS Code to ‘save as UTF-8’ (a mismatch between the encoding declaration and the actual encoding).” This is the problem I ran into first, the one I wrote about in Part 1 of this series.
The second is the intended users. Rev.1 listed three: “Japanese-language users,” “multilingual users,” and “non-engineers.” In Rev.2, I added a fourth: “developers, such as Vibe Coders.” These are the people who, even when an AI prompts them to re-save as UTF-8, find that it does not fix things.
The third is the plan for the next work. I added a feature for rewriting the character-encoding declaration written inside a file, under the number F-14 (a serial number assigned to each feature), and made it the target of the next work. At the same time, I marked this day’s international character encoding support (F-10) as “done.”
The fourth is that I newly created the “How It Works” and “Limitations” sections. In the limitations table, I wrote the two items corresponding to what the tests had caught that day: “GBK and EUC-JP overlap in byte range, so short texts can be mistaken for one another” and “candidates earlier in the trial order take priority.”
518 Test Files
Saved at the same time as the requirements document were the test files (fixtures). Rather than making them by hand, I ran a single generation script that Claude Code wrote, and made them all at once. The script is built only from parts that come with Python from the start.
Here is what’s in it. 14 kinds of character encoding × 3 kinds of text × 6 kinds of extension (.txt .html .xml .css .csv .md), each as one pair: the file before conversion, and the answer file saying what it should become after conversion (.expected). That makes 504. To that, 12 files that are UTF-8 to begin with, and 2 broken files. That comes to 518 in total. The ones paired with an answer file number 252 pairs.
These 518 files will play an active part in the blog posts that follow. I will leave that story for later.
At the End of That Day
The work of this installment was saved at 11:41 and 12:41 on March 8. The tests came to 25, all passing. That is the previous 23, plus one Chinese test and one Cyrillic test.
Coming Up Next
Next time steps away from the timeline a little. Why is it that, as I added to Rev.2’s background, mojibake does not get fixed even when you “save as UTF-8” in VS Code? I will write about the mechanism of character encodings itself.
About Soul Resonant Works
Soul Resonant Works is a solo venture developing seven local AI systems.
Starting from zero programming experience, the development is progressing through collaboration with AI.
🌐 Soul Resonant Works:
→ https://sr-works.net/en/index.html
📝 This blog publishes the entire development process as a serialized journal.
CubePlot (free version available)
CubePlot is the first product from Soul Resonant Works — it turns a CSV into a 3D scatter plot you can rotate and zoom. No install, no sign-up: it’s a single HTML file you open in your browser, and it works offline. CubePlot itself does not send the data you load outside your machine — external network traffic is blocked at the browser level (CSP). Start with the free version.
▶ Product page: https://sr-works.net/en/cubeplot/
▶ Get it (commercial license, USD $39 + tax where applicable): https://soulworks8.gumroad.com/l/pzcij
If you found this article useful, please share it.
Leave a Reply