How Accurate Is Photo Grading? What 714 Handwritten Answers Showed

Marking answers is a small job that eats a surprising amount of time. diglabo has a feature that grades a finished worksheet automatically when you photograph it. Even so, the first thought for anyone using it is probably not "how convenient" but "what if the machine misreads and marks a correct answer wrong?"

I assumed at first that a machine would read digits more accurately than a person does. When I measured it, that turned out not to be true.

I measured how accurate it actually is, so here is what came back.

How I measured it

I scanned worksheets a child had actually worked through in pencil, and for 714 boxes containing a handwritten answer I wrote down what a person reading them by eye saw. That becomes the reference.

Those same 714 boxes can be measured again any number of times. Whenever I change how the reading works, I can compare it against the same yardstick.

Results

Out of 714 boxes, this is how many were misread. The lower three were measured under identical conditions.

  • The old approach (processed entirely on the phone): 152 boxes / 21.29%
  • The current approach (gemini-2.5-flash-lite): 10 boxes / 1.40%
  • gemini-2.5-flash: 7 boxes / 0.98%
  • gemini-3.5-flash-lite: 6 boxes / 0.84%

The old approach did all the reading inside the device. It needed no network connection, but it misread one box in five. That number is why I moved to sending the reading out to an external model.

I'm not using the most accurate one because it costs more than three times as much and takes 1.7 times as long. What that buys is 4 boxes out of 714. diglabo is free, so I decided that difference wasn't worth the money.

Where the machine goes wrong is where people hesitate too

Counting the misreadings one by one, 9 of the current approach's 10 errors were "a 0 read as a 6." The remaining one was "a 4 read as a 9."

That bothered me, so I pulled out the actual paper and looked again. The handwritten 0 there did, if anything, look like a 6. The machine's judgment wasn't wrong — that is simply how the character read.

What I thought was a problem of reading accuracy turned, partway through, into a problem of handwriting.

The surprise: the machine's confidence can't be trusted

For each box, the reading also returns how confident it is. I had assumed we could have a person review only the low-confidence boxes.

When I checked, every single misread box had come back marked "confident." Not one box was flagged as low confidence.

A design that picks out the doubtful boxes by confidence does not work. That is why diglabo lets you correct a reading on the spot. It is built on the assumption that the machine's judgment is not the final word.

What I still haven't fixed

Let me be straight about this.

Sometimes it invents a digit in a box where nothing was written. It reads an eraser mark or a shadow on the paper as a number. Countermeasures have reduced it, but not to zero.

In column arithmetic, the rows shift depending on who is writing. In Division with Remainders and Division (÷ 1-Digit Number), a child who skips writing the intermediate "0" ends up putting the remainder one row higher. The machine looks at a fixed position, so it reads a different box. Where there is more than one correct way to write something, I haven't caught up yet.

A caveat about the measurement

Reading the same image again with the same settings, 4 boxes out of 239 (1.7%) came back different from the previous run. With a generative model, that is unavoidable.

So this measurement can only separate differences larger than a few boxes. Among the numbers above, the gaps between 10, 7 and 6 are close to that margin, so I don't read them as "this one is better." What I can read is the drop from 152 boxes to 10.

What I want to measure next

  • Reading on worksheets in languages other than Japanese (English, Chinese, Korean, Spanish)
  • Whether blank boxes are correctly judged blank
  • How often the readings actually get corrected

I won't put a date on it. Once I have measured, I'll add it here.

Marking is the act of checking whether something is right. If the machine misreads, that doesn't mean your child got it wrong. When a reading looks off, correct it on the spot — that is what the correction is there for.


This article is written by a parent who builds diglabo. I am not a teacher or an education specialist. The general guidance here is based on how the topic is taught in Japanese primary schools, checked against Japan's national curriculum guidelines.

diglabo's worksheets follow the Japanese curriculum. What that means if you are outside Japan