This article was originally written in Japanese and translated into English with AI assistance. Please note that some expressions may carry nuances from the original Japanese.
🇯🇵 日本語版はこちら / Japanese version
Series: utf8conv Development Journal (Part 2)
Last time, I wrote about utf8conv’s first day. It ended at the point where, despite losing half a day to the compatibility problem between uv and tkinter, I’d managed to get at least as far as “an empty window opens, and you can pick files and do a dry run.” This time is the story of the following day — day two.
As the title says, the theme of that day was a design change: “separating the core from the GUI.” Honestly, it’s a little embarrassing for someone with no programming experience to talk about “redoing the design.” Even so, the sense that “if I keep growing the day-one code as it is, I’ll hit a dead end” is something you can vaguely pick up on once you’re actually doing the work, even as a non-engineer. I think this installment is worth writing as a way of bringing that vague sense into sharper focus.
A Bad Feeling Was Mixed into the Day-One Code
What I had at the end of day one was a single file called utf8_converter_gui.py. Inside that one file, everything was bundled together: “a function that detects the character encoding,” “a function that gathers the target files,” “a function that actually does the conversion,” “the part that builds the GUI with tkinter,” and “what happens when a button is pressed.”
A function is a single chunk of work that’s been given a name so it can be called from outside. If you ask for it by name — “detect the encoding of this file” — you get an answer back without needing to know what goes on inside. Think of it like cutting out one step of a recipe and giving that step its own name.
It works. And since it works, I could have kept adding features to it as it was. In fact, on day one, when I asked Claude Code, “Shall we grow this a bit more and add a real conversion mode and a backup feature?”, the answer came back, “Yes, we can.”
But when I reported that exchange to the Claude.ai side (the chat version of Claude), a slightly different perspective came up. “As things stand, you’d have to launch the GUI every single time just to check whether the conversion logic works. Doesn’t that make tests hard to write?”
Now that it had been pointed out, it was exactly right. In the current structure, the only way to check whether a conversion works correctly was for a human to click a button with the mouse. Even if I tried to automate an action like “convert a Shift-JIS file and check the result,” the GUI sitting in the middle meant it couldn’t easily be turned into a script (something like a written procedure that runs automatically from top to bottom, without waiting for anyone to operate it). In other words, it was a structure that didn’t get along well with the whole idea of automated testing.
I’d faintly sensed this “tests are going to be hard to write” feeling myself on day one. But I’d left it alone, thinking, “Tests can wait.” Claude.ai didn’t let that slide, and gave me a gentle push: “If you separate them now, things will be easier later.”
What One Word, “Separation,” Really Means
The structure Claude.ai proposed was simple. Split the file in two. One is core.py, dedicated to logic — everything that “works without opening a screen,” like encoding detection and the conversion process, goes in there. The other stays as utf8_converter_gui.py, and keeps only the part that builds the screen with tkinter and the part that calls core.py when a button is pressed.
According to Claude.ai, this way of thinking — “separating into layers” — is one of the classic basic patterns in software engineering. It also explained that this is essentially the same thing as the three-layer architecture — “GUI layer, business logic layer, data access layer” — that I’d been designing for DataMigrator up to the previous articles. The difference is the order: with DataMigrator, I split things into three layers at the design-document stage, whereas with utf8conv, I wrote everything as one piece first and split it apart afterward.
At that moment, one thing clicked for me. There are probably people who can write things cleanly separated from the very start, but when you write your first GUI app on your own, it inevitably ends up as one lump at first. Then, somewhere around the point where it starts working, you realize “this might not be good enough,” and you end up splitting it again. This redo isn’t a “failure” — if anything, it’s just part of the normal development process. Naturally, the same thing happens in vibe coding too.
Giving Functions Precise Names Again
While doing the separation, there was one more small discussion: about function names.
In the day-one code, the main processing was bundled into a function called detect_and_convert. As the name says, it’s a function that does both “detect” and “convert.”
As the separation work went on, Claude.ai pointed out, “This function name doesn’t make its responsibility clear.” “Detecting” and “converting” are separate responsibilities. For example, in dry run mode you only want to detect. In real mode, you take the detection result and convert as well. A function that does both ends up branching internally, which makes it harder to read. In that case, it’s better to split the function in two and give each its own name.
As a result of that discussion, the functions were organized like this: detect_encoding(path) is a function that only detects the encoding and returns it. run_conversion(...) is a function that actually runs the entire conversion process. The two combine to handle both the dry run and the real conversion.
A “discussion about function names” may sound technical and dull, but for me it was a small discovery. A function name is “a signboard that tells others (or your future self) what that code does,” and if the lettering on the signboard is vague, the work inside becomes vague too. There’s something a little similar here to the feeling of deciding on the title of a stage play or the name of a band’s song. Once the title is settled, the content stops wobbling. I found myself oddly convinced that function names in software had the same effect.
The latin-1 Fallback as a “Catch-All”
While doing the separation, I also sorted out the encoding-detection logic. Since utf8conv follows the zero-additional-packages policy, high-accuracy automatic detection like chardet isn’t available. Instead, it detects with simple logic: “try several encodings in a predetermined order, and adopt the first one that reads successfully.”
The order goes like this. First, it tries whether the file can be read as UTF-8. If it can, that file doesn’t need converting, and that’s the end of it. Only when it can’t does it start trying the candidate list from the top, in order. My initial list was something like “Shift-JIS, CP932, EUC-JP.” If the target was Japanese files, that seemed like plenty.
But here I noticed something: the problem that if it ran into a file that “can’t be read with any of the encodings,” processing would stop. An ordinary text file should almost certainly be readable with one of them, but there’s a chance that, very rarely, a file comes along where “every one of them errors out.”
The solution that came up was the idea of placing an encoding called latin-1 (ISO-8859-1) at the very end of the trial order. latin-1 maps every byte (0x00-0xFF) to some character, so it has the property of “being able to read any byte sequence without raising an error.” In other words, if you keep it there as the very last fallback — a backup measure you prepare for when all the main candidates have failed — you can at least prevent the situation of “processing stopping.”
However, when a file is forced through latin-1, the characters may not be accurate. If a file that was really Japanese gets read as latin-1, a garbled result gets output as “converted.” So latin-1 is strictly “the last-resort fallback for when nothing else could read it,” and in day-to-day use the expectation is that files get read by the Japanese encodings earlier in the list.
This, too, was something where I only thought “I see” once it was pointed out to me. The idea of “deliberately putting a weak rule at the end” simply wasn’t in me. I came to vaguely understand that design isn’t just lining up the strong parts — it’s work that includes deciding where to place the weak parts, too.
★ That said, let me say up front: this decision gets overturned later. Keeping latin-1 there makes the state of “nothing could read it” itself disappear, which creates a different problem — the tool silently converts even files that aren’t text. Today’s utf8conv doesn’t include latin-1 as a candidate. I’ll write about how that came about in a later installment.
The [build-system] Section — an Unglamorous Trap
Let me write about one more pitfall I ran into that day. Technically it’s a very small thing, but it left an impression on me.
After separating the core from the GUI, at some point that day, I got an error where import utf8conv failed.
★ To be honest, I couldn’t pin down whether this happened right after the separation work or a bit later that same day. When I cross-checked the records from that time while writing this article, two records said different things (the work log said it was during the later work; the summary said it was during the separation work). That day, I ran three sessions all within the same March 8, and the boundaries between them have blurred. Either way, it’s something that happened that day and got resolved that day.
I’ll start from digging into the cause.
First, there’s a file called pyproject.toml. It’s a file that gathers a project’s settings in one place — the project’s name, its version, the parts it needs, and so on. And uv (pronounced “you-vee”) is the tool that also appeared last time. It prepares Python itself and the parts you need, and runs your program.
The single line import utf8conv means “load the bundle named utf8conv.” This “bundle” is called a package.
The problem was that utf8conv’s pyproject.toml had no section with the heading [build-system]. Roughly speaking, this section is a statement that says, “This folder can be assembled into a distributable part.” uv uses whether this statement is present to decide “whether it’s okay to treat this as a package.” If there’s no statement, it considers “this isn’t a package,” so import utf8conv can’t find where to go, and you get an error.
What’s interesting is that I had never once run into this problem on the DataMigrator side. When I later lined up the two pyproject.toml files side by side, DataMigrator didn’t have [build-system] written in it either.
Instead, in the test settings, there was just one single line written like this:
pythonpath = ["src"]
It means “the main body of the program is inside a folder called src.” Before starting the tests, the tool that runs them (pytest) builds a list of “places to look for the main body” (the search path). This one line adds the src folder to that list. Then import utf8conv ends up in a state where it gets found even without being stated as a package, because the tool has been told directly where to look.
Either state that it’s a package and let it be found, or tell the tool directly where to look. Either way, the test’s import goes through. However, the latter (pythonpath) only works while the tests are running — it has no effect when people use what you’ve distributed. DataMigrator happened to get through with the latter, and utf8conv had neither.
In other words, I had simply been sidestepping the same problem by a different method, by chance. Even though utf8conv was started using DataMigrator as a template, it hadn’t inherited that one line. And I hadn’t noticed on my own that it hadn’t been inherited.
The reason this was a slightly painful experience is that what tripped me up was the assumption that “it worked in DataMigrator, so it should be fine.” Carry the success of your first project straight into your second, and you won’t notice the differences in the settings files. Looking back, “the second project leans too heavily on the success of the first” may be a lesson that applies to anything, not just vibe coding.
The End of Day Two
By the end of that day, the core and the GUI had been separated, the conversion logic had proper tests, the latin-1 fallback was built in, and the [build-system] trap had been dealt with. 17 tests, all passing. Test coverage for core.py: 95%. I still don’t properly understand what the numbers mean, but I do at least understand that “apparently it’s high.”
Compared to day one, as code it was an unglamorous kind of progress. No new features had been added, and from the user’s point of view, nothing had changed. Even so, the internal structure had changed a great deal, and from here on it had moved into a state where “features are easy to add,” “tests are easy to write,” and “bugs are easy to find.”
Day one was “build something that works”; day two was “reshape it into a form that’s easy to grow.” Thinking back, I believe that order was just right. If I’d tried to write it in an “easy-to-grow form” from the start, I probably wouldn’t have had anything working by the end of day one — and even if I had, I’d surely have been working without any confidence that “this is the right way.”
Coming Up Next
Next time (UC03) is utf8conv’s third day. In terms of the date, it’s actually the same March 8 as this one. This is the day the MVP (minimum viable product) finally gets completed. The themes are: showing the progress of the conversion on screen, writing integration tests to “check the whole flow without launching the GUI,” and the episode of what happened when I ran a dry run on my own real data (the SRW folder) for the first time.
About Soul Resonant Works
Soul Resonant Works is a solo venture developing seven local AI systems.
Starting from zero programming experience, the development is progressing through collaboration with AI.
🌐 Soul Resonant Works:
→ https://sr-works.net/en/index.html
📝 This blog publishes the entire development process as a serialized journal.
CubePlot (free version available)
CubePlot is the first product from Soul Resonant Works — it turns a CSV into a 3D scatter plot you can rotate and zoom. No install, no sign-up: it’s a single HTML file you open in your browser, and it works offline. CubePlot itself does not send the data you load outside your machine — external network traffic is blocked at the browser level (CSP). Start with the free version.
▶ Product page: https://sr-works.net/en/cubeplot/
▶ Get it (commercial license, USD $39 + tax where applicable): https://soulworks8.gumroad.com/l/pzcij
If you found this article useful, please share it.
Leave a Reply