This article was originally written in Japanese and translated into English with AI assistance. Please note that some expressions may carry nuances from the original Japanese.
🇯🇵 日本語版はこちら / Japanese version
Series: utf8conv Development Journal (Part 1)
Starting with this installment, I’m beginning the development journal for another product that grew up alongside DataMigrator: “utf8conv” (UC for short).
utf8conv is a very simple GUI tool that does nothing more than batch-converting files saved in a non-UTF-8 encoding into UTF-8. The name is exactly what it sounds like: “utf-8 converter.” Unlike DataMigrator, there’s no lengthy design document behind it, and in scale the two aren’t remotely comparable — utf8conv is much smaller.
Still, being small didn’t mean it was easy. Right from day one I ran into a compatibility problem between uv and tkinter that ate half a day, and partway through I struggled under a “zero additional packages” constraint I’d set for myself — there were a fair number of stumbles along the way.
This first installment starts with the story of why I decided to build this tool in the first place.
While Talking with Claude, Mojibake Was Quietly Happening
Around February 2026, alongside developing DataMigrator, I was also doing other work while chatting with Claude — tweaking HTML files here and there, tidying up old documents, that kind of miscellaneous task.
One day, in the middle of asking Claude for advice on something, I found a garbled file on my PC. Opening it, the Japanese parts had turned into something like “譁?ュ怜喧縺” — a string of symbols that unmistakably screamed “the character encoding is broken.” Having used a Mac for 20 years, this kind of mojibake is a familiar sight I’d run into more than a few times.
Normally I’d have shrugged it off with an “ugh, again” and dealt with it however. But this time, since I happened to be in the middle of migrating data from my old iMac to a MacBook Pro, and had Claude on hand to ask, I tried “How do I fix this file?” Claude’s answer was straight out of the textbook: “The original file is probably saved in a different encoding, like Shift-JIS. Please re-save it as UTF-8 in VS Code.”
So I did as told: opened it in VS Code, chose “Save with Encoding” → UTF-8, saved it, and reopened it in the browser.
—It didn’t fix it.
On screen, the mojibake was still dancing around, unchanged. VS Code’s own status bar at the bottom said “UTF-8,” so the file was surely UTF-8 now. When I told Claude, “I saved it but it’s still not fixed,” Claude thought for a moment and suggested, “Try explicitly specifying the encoding again when you reopen and save it.” I tried that too. Still no fix.
At this point, doubts were starting to creep in: “Is this a VS Code bug?” “Is the AI telling me something wrong?” Doubting one or the other is one thing, but once you start doubting both at once, you lose your footing entirely.
So, What Is a Character Encoding, Anyway?
Let me back up for a moment and lay out some groundwork.
Inside a computer, every character is stored as numbers — a sequence of bytes. The rule for how to read that byte sequence and map it to characters is what we call a “character encoding.” There isn’t just one such rule; many have been created over the years, for different languages and eras — Shift-JIS or EUC-JP for Japanese, GBK for Chinese, CP1252 for Western European languages, and today’s standard, UTF-8.
The same character will produce a different byte sequence depending on which rule you use. Put the other way around: read the same byte sequence under a different rule, and you get something completely different. That’s mojibake.
Let’s actually try it. Take a string saved under one language’s character encoding, then read it as if it were a different language’s encoding, and here’s what you get:
Below, a Japanese file is read as Chinese (GBK), a Chinese file as Japanese (EUC-JP), and a German file as Cyrillic (CP1251).
Japanese こんにちは 文字化け → 偙傫偵偪偼 暥帤壔偗
Chinese 你好 世界 编码 → 低挫 弊順 園鷹
German Grüße Straße Fußball → GrьЯe StraЯe FuЯball
Look at the German example. Grüße has become GrьЯe. The alphabetic letters survive intact; only the characters carrying an umlaut or an eszett break, turning into Cyrillic letters. Unlike the Japanese case, where the whole string turns unreadable, here only part of it corrupts. Exactly which parts the two rules share, and which parts differ, shows up directly in the result.
And the mojibake that opened this article runs in the opposite direction. Take something saved in UTF-8 and read it as if it were Shift-JIS, and you get this:
Japanese "mojibake" (文字化け) → 譁蟄怜喧縺
In both directions, the file itself isn’t broken. The byte sequence is intact, safely saved exactly as it was. What’s broken is the “how to read it.”
The Cause: A Mismatch Between the Byte Sequence and the Declaration
Digging into the cause together with Claude, here’s what it eventually turned out to be.
The problem file was an HTML file. HTML files carry a declaration line near the top of the file that says, in effect, “this file is written in this encoding” — specifically, something like <meta charset="Shift_JIS">. The browser sees this declaration and decides, “ah, then I’ll read it as Shift-JIS.”
When you choose “Save with Encoding: UTF-8” in VS Code, the file’s contents — the actual byte sequence — really are converted to UTF-8. VS Code isn’t lying about that part. But the line that says <meta charset="Shift_JIS"> is left exactly as it was.
That’s because an editor’s “save with a different encoding” feature converts the file’s byte sequence — it doesn’t read the document’s contents and rewrite them. So the string <meta charset="Shift_JIS"> just gets stored, unchanged, inside a UTF-8-encoded file.
And I don’t think that’s careless of the editor — I think it’s the correct behavior. An editor that took it upon itself to rewrite the body text because “this declaration doesn’t seem to match reality, so I’ll fix it for you” would be far more alarming.
Still, the upshot is that the file is left holding nothing but a contradiction.
The result is a mismatch like this: the file’s actual contents are UTF-8. But the top of the file declares “this is Shift-JIS.” The browser trusts the declaration first. Believing it’s Shift-JIS, it tries to interpret the UTF-8 byte sequence as Shift-JIS. Naturally, the result isn’t meaningful Japanese. That result is the mojibake from the start of this article.
The moment I heard this explanation, two thoughts rose up in me at once: “I see,” and “wait, couldn’t this be turned into a tool?” The problem I’d run into was probably one that plenty of VS Code users out there, somewhere, had also run into — and given that it even included the situation of “an AI kindly walking you through the wrong fix,” it seemed like something with real reproducibility. In particular, since I was already building DataMigrator, a product for organizing and migrating data files, with the underlying premise of getting file migration between my own Macs properly sorted out, I was going to need to correctly interpret the encodings of a large, disorganized pile of files anyway. That’s also where requirements like being able to do a dry run in advance, actually changing the encoding properly, handling a large number of files at once, and having a safety net to revert if needed, all found their way into the spec.
To begin with, a tool that just batch-converts files from non-UTF-8 to UTF-8 is the kind of thing that could reasonably have existed for a long time already. But one that’s a GUI, needs no extra installation, can safely do a dry run (check only), and correctly rewrites HTML/XML declarations too — all of that together — I couldn’t find, within the range I searched.
By the way, this problem — mojibake caused by a mismatch between the declaration and the actual encoding — wasn’t limited to browsers. The same thing happens in the Quick Look preview you get on macOS Finder by selecting a file and pressing the space bar. I confirmed this later by making four test files.
A file whose contents are UTF-8 but whose declaration says “Shift-JIS” — garbled, even in Quick Look. A file where both the contents and the declaration are UTF-8 — readable. One where both are Shift-JIS — also readable. And a file with the declaration removed entirely — readable, correctly.
In other words, Quick Look can actually work out the right answer just by looking at the contents. Yet even though it can work it out, when a declaration is present, it trusts the declaration instead and breaks. It displays correctly only when there’s no declaration at all — a slightly perverse result.
From here on, this is less about character encoding itself and more about how the software is built. Software that handles character encodings is, for the most part, built to “auto-detect by default, but let you choose manually if needed.” Browsers work this way too. But the lightweight preview feature built into the OS — the one that doesn’t even count as launching an app, that opens instantly when you select a file — doesn’t seem to have that manual entry point anywhere. Maybe it exists somewhere if I looked hard enough. But the act of hunting for it is a hassle in itself, and even if I found the setting and changed it, some other encoding would stop displaying correctly instead — which would defeat the purpose.
Rather than fiddling with the display side to fix it, fix the file itself. That, in the end, is what utf8conv is trying to do.
UC’s Direction
And so, right as DataMigrator’s M1 wrapped up, another small project — “utf8conv” — got underway.
I wrote a requirements document for it too, but it never grew past 30-plus chapters the way DataMigrator’s did. It’s a modest requirements document, running to just a few chapters at most. Even so, I settled a few important policies right at the start.
The first was “zero additional packages” — a constraint that everything had to be done using nothing but Python’s standard library. I initially considered using a well-known library called chardet for encoding detection, but I judged that skipping chardet would buy me benefits like “it runs the moment it’s installed,” “no dependency headaches,” and “distribution stays simple.” Detection accuracy might suffer a little as a result, but I figured I’d make up for that with the “try encodings in order from the top” approach described later.
The second was “dry run first.” By default, the tool doesn’t convert anything — it just runs in a mode that logs what would happen if you converted, and only performs an actual conversion once the user explicitly chooses “really convert.” Just like DataMigrator’s “never delete” philosophy, this leans toward the safe side. Since undoing a character-encoding conversion once it’s done is a hassle, I judged that a confirm-first mechanism was indispensable.
The third was “automatic backups.” When an actual conversion runs, the pre-conversion file is automatically moved into a backup folder. Even if a user chooses to skip the dry run and go straight to a real conversion, having a backup means they can still roll back — a last line of defense.
Written out, these three policies come to only three lines, but they ended up defining utf8conv’s character down the line.
On Day One, uv and tkinter Get Into a Fight
Once the requirements document was done, I started actually moving my hands in Claude Code. Same as with DataMigrator: create the project’s box, spin up a Python environment, get a minimal GUI showing — that route. For the GUI, in keeping with the zero-additional-packages policy, I decided to use tkinter from Python’s standard library.
But when I tried to run it, the window wouldn’t open. The error wasn’t particularly helpful either — all I got back was a message to the effect of “Tcl/Tk not found.”
Having Claude Code dig into it, the cause became clear. uv (the Python environment manager I’d adopted during DataMigrator’s M1) ships its own Python binaries, and those binaries don’t bundle Tcl/Tk — the GUI library tkinter relies on internally. The Python that uv distributes is, strictly speaking, “just the Python interpreter” — it doesn’t include GUI-related extra runtimes. That’s a reasonable design choice on uv’s part, but it’s a trap if you’re trying to use tkinter.
The fix was to install Python and python-tk via Homebrew, then rebuild uv’s virtual environment on top of that Homebrew Python. Written as commands, the flow looks roughly like this:
brew install python@3.12 && brew install python-tk@3.12
uv venv --python /opt/homebrew/bin/python3.12
Written down, it’s only a couple of lines, but it took about half a day of trial and error to pin down the cause. Neither uv nor tkinter is at fault — it’s the specific combination that’s unusual, the kind of problem that’s genuinely hard to reach on a first encounter.
Still, this experience paid off later. From then on, every Python project with a GUI — both utf8conv and DataMigrator — got standardized on this “uv virtual environment built on Homebrew Python” setup, and I never ran into the same problem again. One more piece of experience banked, so I wouldn’t hit the same wall twice.
The End of Day One
By the end of day one, a directory structure suited to software development was in place, and the source file utf8_converter_gui.py was up and running.
Launching utf8_converter_gui.py opened a tkinter window. The conversion logic itself was already written by this point. I could already select which file extensions to target for conversion, and there was already an “add custom extension” input field where I could add my own. On top of that, the GitHub Actions test automation was running, and the first push to the private repository was done.
That said, the current source code has the conversion logic and the GUI mixed together in a single file. In this shape, there’s no way to pull out just the conversion logic and verify it on its own. In fact, the only test I managed to write that day was a single one checking whether the version number could be read correctly. Splitting these apart so each piece could be tested individually — that became the next day’s homework.
Oddly enough, I didn’t get the same heavy sense of “taking the first big step into a large development effort” that I’d felt when M1 wrapped up on DataMigrator. utf8conv’s first day felt lighter than that — more like an oddly optimistic “maybe this’ll be done in three days.” Maybe it’s because the smaller scope let me picture the finished product right away.
And in fact, utf8conv’s development would go on to produce a working MVP within a genuine handful of days after this. Three days including a weekend, six chat sessions (U1 through U6), and all the features were implemented — that was the pace. Along the way, I went through a few design decisions that pushed a bit further than typical vibe coding, like “separate the core from the GUI,” “put the latin-1 fallback last,” and “write tests as data instead of code.” I’ll get into those starting next time.
Coming Up Next
Next time is the story of utf8conv’s second day (session U2). As of day one, the conversion logic and the GUI were still mixed together in one file. Since that makes it impossible to write tests (you can’t check the logic without launching the GUI), I end up separating the core (logic) from the GUI (screen). Changing a design for the sake of “so we can write tests” was, for me, a first. Along the way, I also get into an idea called the “latin-1 fallback,” for implementing character-encoding detection using nothing but Python’s standard library.
About Soul Resonant Works
Soul Resonant Works is a solo venture developing seven local AI systems.
Starting from zero programming experience, the development is progressing through collaboration with AI.
🌐 Soul Resonant Works:
→ https://sr-works.net/en/index.html
📝 This blog publishes the entire development process as a serialized journal.
CubePlot (free version available)
CubePlot is the first product from Soul Resonant Works — it turns a CSV into a 3D scatter plot you can rotate and zoom. No install, no sign-up: it’s a single HTML file you open in your browser, and it works offline. CubePlot itself does not send the data you load outside your machine — external network traffic is blocked at the browser level (CSP). Start with the free version.
▶ Product page: https://sr-works.net/en/cubeplot/
▶ Get it (commercial license, USD $39 + tax where applicable): https://soulworks8.gumroad.com/l/pzcij
If you found this article useful, please share it.
Leave a Reply