A UPC check digit catches every one-digit error, but not every swap

016345678903 is a valid synthetic UPC. Change any one of its twelve digits and the number fails its arithmetic check. Swap the neighboring 1 and 6, though, and 061345678903 passes.

The barcode has not suddenly decided that two products are the same. A scanner has other structural clues, and a store still has to find the number in its product database. But the example exposes the modest promise made by the little digit at the far right of a UPC: it is very good at catching one wrong digit, not all possible errors.

That last digit is called the check digit. GS1 describes it as an integrity check calculated from all the digits before it. The recipe is pleasantly small. Starting from the right side of the payload, multiply alternating digits by 3 and 1, add the results, then choose a final digit that brings the total up to a multiple of ten. GS1 publishes the same method for calculating it by hand.

For the familiar teaching example 01234567890, the weighted sum is 85. The check digit is therefore 5, producing 012345678905. A scanner can repeat that arithmetic before trusting what it read.

What the tiny test caught

I ran a bounded exhaustive test on that synthetic 12-digit number. There are 12 positions and nine wrong digits that can replace the correct digit in each position, so the test made 108 single-digit substitutions. The check rejected all 108.

synthetic UPC: 012345678905
single-digit substitutions tested: 108
undetected: 0

The result does not depend on that particular number. A one-digit change alters the weighted total by the change itself or three times the change. Neither can be a multiple of ten unless the digit did not change at all.

Swaps are trickier. I also tested every ordered pair of unequal decimal digits as neighbors: 90 possibilities. The check caught 80 and missed 10. Every miss exchanged digits five apart: 0 with 5, 1 with 6, and so on. That is why the opening pair survives. Moving the 1 into the 6's weighted position and the 6 into the 1's changes the total by a multiple of ten, so the check digit has nothing to complain about.

This does not mean UPCs have an alarming one-in-nine failure rate. The test deliberately examined one narrow kind of manufactured error, ignored the barcode's physical encoding, and gave every possible neighboring pair equal weight. It tells us about the check digit's shape, not the real-world error rate at a checkout.

The barcode began with a different shape

There is a nice anniversary hiding behind this arithmetic. On October 7, 1952, Norman Woodland and Bernard Silver received U.S. Patent 2,612,994 for a machine-readable classification system. Their patent showed both straight lines and a version bent into concentric circles so a package would not need to be carefully oriented before scanning.

The bull's-eye was clever and difficult to print reliably. Two decades later, George Laurer's IBM team developed the rectangular Universal Product Code, whose vertical bars handled printing better while scanners dealt with orientation. IBM's history of the UPC says the final symbol uses 30 black bars and was formally adopted in 1973.

The visible bars did the dramatic work, translating a number into light and dark. The last digit took the duller assignment: notice when something probably went wrong. It cannot understand the product or repair a damaged label. It just multiplies by 3, multiplies by 1, and refuses every single-digit mistake. For one digit of overhead, that is a respectable little job.