Completing the MVP — When Detection Cannot Be Perfect, Show the Work and Let a Human Check

This article was originally written in Japanese and translated into English with AI assistance. Please note that some expressions may carry nuances from the original Japanese.

🇯🇵 日本語版はこちら / Japanese version


Series: utf8conv Development Journal (Part 3)


Last time, we got as far as splitting utf8conv’s conversion internals (the core) and its screen (the GUI) into separate files, so that tests would be easier to write. This time picks up from there. The date is the same as last time: March 8 (a Sunday). In this series I call it “day three,” but strictly speaking, it was the third round of work (session) carried out on that same day.

If I had to give day three’s work a title, it would be “GUI completion, integration tests, MVP finishing touches.” MVP is an acronym for “Minimum Viable Product,” and it means “the smallest thing that’s usable for now.”

Let me state up front the one thing I want to say in this installment. On that day, utf8conv could not make its character-encoding detection perfect. Instead, I made what it was doing visible on screen, and shaped it so that a person could check before anything got rewritten.

That is what the MVP consisted of that day.

Homework from Last Time — Showing Progress on Screen

The first item on the “next things to do” list I left at the end of the previous session was “brushing up the GUI (progress bar, status improvements, etc.).” That’s where I started on this day.

I did three things.

First, I put the progress bar (a horizontal bar that shows how far along things are) on the same row as the run button. Second, I moved the one line that reports the current state (the status bar) to the very bottom of the window, and during processing I made it show which item out of how many it’s on — for example, “Processing… 3/10 items.” Third, I finished the color coding of the log. Converted files are green, dry run (a trial run that doesn’t actually rewrite anything and only checks what would be converted) is blue, errors are red, and headings are yellow. I also prepared gray for skips. However, skipped files weren’t logged one by one; the design was to output only a count at the end, like “UTF-8 (skipped): ○ items,” so a gray line never actually appeared on screen.

All three were work on the screen side; I didn’t change the way conversion itself works. It was work to get “what the conversion is doing right now” in front of the user’s eyes.

Only the One in Charge of the Screen Touches the Screen

Let me take a moment here to talk about how things are put together.

utf8conv runs its conversion processing on a “separate thread.” A thread is a flow of work that proceeds side by side with others inside a program. If there’s only one flow, then while the conversion is running, redrawing the screen and responding to buttons have to line up behind it and wait. So the conversion is handed off to a separate flow, and the screen’s flow is kept free.

This structure — the screen part and the conversion processing on separate threads — wasn’t something added on this day. It was already that way in the code Claude wrote on day one (March 7).

The tool that builds the screen is tkinter (pronounced “tee-kay-inter”; a collection of parts for building screens that comes with Python from the start). On this day, I decided to leave progress-bar updates to the main flow as well. This is so that tkinter can be used safely from more than one flow. The conversion flow doesn’t touch the screen directly; instead, using the form self.after(0, ...), it asks the main flow, “Please display this.” The side that was asked carries it out when it has a free moment.

This way of asking was also already in the day-one code, in two places: when outputting a line of the log, and when announcing that everything had finished. On this day, I added one more — progress-bar updates — using the same way of asking.

Along with that, I also added one opening on the conversion-internals side (core.py) for reporting progress. Because I made this opening something “you can use or not use,” not a single one of the tests I wrote last time had to be changed.

Tests as a Full Run-Through

On this day, I added five “integration tests.” Whereas last time’s unit tests checked the parts one at a time, integration tests check a whole sequence with the parts connected together. In theater terms, it’s the equivalent of a full run-through rehearsal.

The five tests cover the following five things:

  1. A Shift-JIS file gets converted to UTF-8
  2. In a dry run, files are not rewritten
  3. A backup gets created
  4. Files that are already UTF-8 get skipped
  5. And a folder with mixed character encodings can be processed all the way through

In every case, pytest (a tool that runs tests automatically) checks things by calling the conversion-internals functions directly, without opening the screen. Because I’d separated the internals from the screen last time, it had become possible to test the whole flow without launching the screen.

With this fifth test, I stumbled once.

A short piece of text written in EUC-JP was detected as Shift-JIS. utf8conv at the time was built to try candidate encodings in a fixed order and take the first one that could read the file as the answer. In that order, Shift-JIS comes before EUC-JP. A short EUC-JP text can sometimes be read as Shift-JIS as well.

What I fixed at that point was not the detection, but the test. I changed the result the test demands from “which encoding it was detected as” to “the converted file can be read as UTF-8,” and got it to pass that way.

Where to Catch What Cannot Be Fully Fixed

About this stumble, the work that day sorted things out as follows. Misdetections like this are inherently unavoidable unless you use an additional part that detects statistically. And since we have decided not to use additional parts, the flow in which the user checks the results with a dry run becomes important.

From the very beginning, utf8conv had a rule that it would run using only the parts that come with Python out of the box (I called this “zero additional packages”). Within that rule, detection cannot be made perfect. So on this day, I decided to put a place for checking not on the detection side, but on the user’s side. In the instruction manual (README) I wrote that day, I also included “Dry run first — check only by default. Actual conversion only when explicitly instructed.”

Reading it back now, that day’s work looks connected by a single line. Show the progress, color the log, and make what’s happening visible on screen. Before rewriting anything, have a person check with a dry run. In other words, on the premise that detection can be wrong, I prepared a place where mistakes can be noticed.

Running a Dry Run on My Own Folder

In this session, I also tried it on a real folder of my own. The target was the SRW (Soul Resonant Works) folder, ~/Documents/SRW, which contained 51 files. Running it as a dry run, all 51 files were detected as UTF-8, and none needed converting. Since it was a dry run, nothing was written to the files.

Finishing Touches — CI and README

Finally, there were two small finishing touches.

One was revisiting the CI settings. CI (continuous integration) is a mechanism that automatically runs the tests every time you send code to GitHub, the place where the whole set of code is kept (the repository); here, it uses a service called GitHub Actions. This had been set up on day one, and at that time it was configured to run the tests on two versions of Python: 3.8 and 3.12. In the previous session, I had aligned the whole project on Python 3.12, so on this day I narrowed CI down to 3.12 only.

The other was writing a new README (the instruction manual placed at the very top of the repository). It sums up on a single page what the tool does, which character encodings it supports, and how to launch and use it.

At the End of That Day

The work of this session was saved at 8:49 on March 8. The previous session’s work had been saved at 8:30.

At this point, there were 23 tests in total, all passing. The breakdown: 17 unit tests from last time, 1 version check from day one, and 5 integration tests from this day. Of the conversion internals (core.py), the proportion of lines that actually ran during the tests (coverage) was 95%.

utf8conv at the end of this day was in this state: you pick a folder, check with a dry run what would be converted, and when you actually convert, the original files remain as backups, and you can watch it all happen through the progress bar and the color-coded log. There were still things it didn’t have. The candidate encodings to try were just five: four Japanese-related ones and the final catch-all (latin-1). There was no saving of the log, and no feature for picking just one file and converting it.

Coming Up Next

Next time continues the same March 8. It’s the story of adding 11 more candidate encodings — Chinese, Korean, Cyrillic, Eastern European, and Western European — for a total of 16. The story of using a tool called PyInstaller to export utf8conv in the form of a Mac app (.app). And the story of rewriting the requirements document as Rev.2 and automatically generating 518 test files all at once.


About Soul Resonant Works

Soul Resonant Works is a solo venture developing seven local AI systems.
Starting from zero programming experience, the development is progressing through collaboration with AI.

🌐 Soul Resonant Works:
→ https://sr-works.net/en/index.html

📝 This blog publishes the entire development process as a serialized journal.


CubePlot (free version available)

CubePlot is the first product from Soul Resonant Works — it turns a CSV into a 3D scatter plot you can rotate and zoom. No install, no sign-up: it’s a single HTML file you open in your browser, and it works offline. CubePlot itself does not send the data you load outside your machine — external network traffic is blocked at the browser level (CSP). Start with the free version.

▶ Product page: https://sr-works.net/en/cubeplot/
▶ Get it (commercial license, USD $39 + tax where applicable): https://soulworks8.gumroad.com/l/pzcij


If you found this article useful, please share it.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *