One of my Rails apps logs into a third-party portal on a schedule. That portal puts a captcha in front of every login: a 290x80 PNG, six glyphs, drawn over a fixed background texture.
A human types it. A cron job cannot. So the controller has two paths: create takes whatever a person typed, and auto_create hands the image to an offline solver. Both open the session the same way; only the second one is allowed to fail with a “needs a human” error and fall back to the screen.
The interesting part is not the HTTP. Reading six glyphs off a background is four image operations, and all four are one-line libvips calls. To find out whether that holds up, I wrote a captcha generator with rubyvips that produces images whose answers I already know, then pointed the solver at it.

A generated captcha. Tilt, per-glyph scale, uneven paper, and noise are deliberate: each one exists to break a naive reader
Versions: Ruby 4.0.7, ruby-vips 2.3.0, libvips 8.18.7. The solver (segment, normalise, classify, abstain) is 189 lines of Ruby, and no gem beyond ruby-vips.
Part 1: creating a captcha you already know the answer to
The generator has one job beyond drawing digits: the background must be byte-identical across every sample. Background subtraction only works if there is a background to subtract, so the paper is built once and memoized.
# sample.rb
class Rng
def initialize(seed) = @s = seed
def next_f = @s = (@s * 1_103_515_245 + 12_345) & 0x7fffffff
def rand = next_f / 0x7fffffff.to_f
def range(a, b) = a + rand * (b - a)
def int(a, b) = a + (rand * (b - a + 1)).floor
end
A hand-rolled LCG, not Random, so every number in this post is reproducible from a seed. eval.rb runs 300 samples from seeds 50000–50299; you get the same table on any machine.
The paper has to be ugly in specific ways:
# Uneven lighting: a vertical tone drift plus a horizontal one, so that
# no single cut point separates ink from paper across the whole frame.
v = 190.0 + (y / H.to_f * 30 - 15) + (x / W.to_f * 16 - 8)
blobs.each do |bx, by, rw, rh|
next if ((x - bx) / rw.to_f)**2 + ((y - by) / rh.to_f)**2 > 1.0
v -= 46
end
That is an 84-level spread across one frame (129 in the middle of a blob, 213 in the bright corner), and it is the single most important property of the generator. I will come back to it in the results.
Digits come from a 5x7 bitmap font. Each one is stamped with independent horizontal and vertical scale, a tilt of ±6°, and its own ink density:
angle = rng.range(-MAX_TILT, MAX_TILT) * Math::PI / 180
box_w = (cell_w * cos.abs + cell_h * sin.abs).ceil
box_h = (cell_w * sin.abs + cell_h * cos.abs).ceil
box_h.times do |dy|
box_w.times do |dx|
# Every destination pixel is inverse-mapped into the font cell, so the
# tilt costs a nearest-neighbour lookup and no resampling buffer.
source_x = (ox * cos + oy * sin) + cell_w / 2.0
source_y = (-ox * sin + oy * cos) + cell_h / 2.0
next unless source_x.between?(0, cell_w) && source_y.between?(0, cell_h)
next if bits[(source_y / vscale).floor][(source_x / hscale).floor] == "0"
pixels[py][px] = (ink + rng.int(-8, 8)).clamp(0, 255)
end
end
Two details are load-bearing and were not obvious:
Ink is chosen relative to the paper, not absolutely. Measure the lightest paper pixel in the digit band first, then put ink 30–170 levels below it. With an absolute ink value, a “dark” sample renders digits that vanish over the bright half of the frame. The sample is not hard; it is impossible.
The gap between digits is wider than the tilt. x += box_w + rng.int(9, 15). A tilted cell is wider than its upright version. When I first used a 4–8 pixel gap, two digits touched and column-projection segmentation fused them into one blob. That belongs in the technique’s contract, not in the generator: projection segmentation assumes a blank column between glyphs. Real captchas that let characters collide need connected-component labelling or a split heuristic instead.
Finally, ±14 of sensor noise over the whole frame, and the array goes into a Vips::Image. That is where the generator stops and the solver begins.
Part 2: reading it back
Stage 1: subtract the background, then threshold
# pipeline.rb
def initialize(image, reference: Sample.paper)
@grey = image.bands == 1 ? image : image.extract_band(0, n: 3).colourspace("b-w")
@difference = reference.subtract(@grey).abs # libvips has no `difference`
@mask = @difference.relational_const(:more, THRESHOLD)
end
Subtract-then-abs is the whole trick. Ink sitting 40 levels below paper in a shadowed corner is now worth exactly as much as ink sitting 40 levels below paper in the light. The lighting gradient cancels, and one threshold works everywhere.
THRESHOLD is 30, and it is not a taste decision. The reference is noise-free and the sample carries ±14, so a noise-only pixel can differ by 14 and the cut has to clear 28. Anything differing by more than 30 is ink.
Stage 2: find the glyphs with a column projection
# Mean ink per column as a single row of pixels: reduce(1, h-1) block-averages
# each column down to one value, which is the projection profile.
def column_profile = @mask.reduce(1, @mask.height - 1)
def ink_columns(min_pixels = 2)
column_profile.relational_const(:more, 255.0 * min_pixels / @mask.height)
end
Contiguous runs of inked columns are glyphs: one pass over a single row of numbers. The threshold is written as “at least two inked rows per column” rather than as a magic number, so the rule survives a different cut.
I raised it from one pixel while chasing a bug where two digits fused. The bug turned out to be geometric (the tilt eating the gap), not the pixel count. It stayed because naming the assumption is worth more than the default.
At THRESHOLD = 30 with ±14 of noise, a noise-only pixel cannot ink a column at all, so this guard does nothing on my samples. It matters for the real solver, where JPEG ringing around a glyph edge does produce isolated pixels.
Stage 3: normalise, so glyphs can be compared at all
Raw crops differ in size, position, stroke thickness and tilt. Comparing them directly compares noise:
# pipeline.rb
def self.call(mask, frame: mask)
return nil if mask.width < 3 || mask.height < 2
left, top, width, height = mask.find_trim(background: 0)
return nil if width < 3 || height < 5
return nil if width > frame.width * 0.6 || height > frame.height * 0.9
mask.extract_area(left, top, width, height)
.resize(WIDTH.to_f / width, vscale: HEIGHT.to_f / height, kernel: "nearest")
.relational_const(:more, 127)
end
Trim to the ink bounding box, resample to a fixed 20x28 canvas, re-binarise. After this, a 5x7 digit scaled 5x lands in the same box as one scaled 7x.
Stage 4: match against averaged templates
One template per digit, built by harvesting glyphs out of 80 generated captchas (480 glyphs, about 48 per digit) and taking the pixel-wise mean:
seen[char] += 1
Pix.rows(normalised).each_with_index do |row, y|
row.each_with_index { |v, x| sums[char][y][x] += v }
end
A single render is an exact match to itself and nothing else. The mean of ~48 of them is a soft, centred average shape that most noisy glyphs still resemble. That is why the templates below have grey edges.

The ten templates. Grey edges are the point: each is the average of roughly 48 tilted, noisy, differently scaled renders
Matching is one libvips call, and the runner-up distance is the confidence:
# Mean absolute difference, done as a libvips subtract-abs-average.
# 0 is identical, 255 is maximally different.
def self.distance(glyph, template) = glyph.subtract(template).abs.avg
def match(glyph)
ranked = @templates.map { |char, template| [ char, Template.distance(glyph, template) ] }.sort_by(&:last)
(best, distance), (_, second) = ranked[0], ranked[1]
[ best, distance, second - distance ]
end

Paper, captcha, difference amplified 3x, and the binary mask. The gradient and blobs vanish in the difference panel: that is the entire argument for this stage

Six segmented glyphs after trimming and resampling. Different sizes, different tilts, one shape space
Four libvips traps I hit while writing this
Every one of these produced a wrong answer or a crash while looking like correct code.
| Call | What bit me |
|---|---|
find_trim |
Without an explicit background: it estimates the background from the image. On a binary mask that estimate is always ambiguous, so it returned the full box every time and nothing was ever trimmed. find_trim(background: 0) is not optional. |
find_trim |
It estimates the background with a percentile filter over a fixed window (libvips’ rank operation). On a strip 2 pixels wide that window does not fit and it raises rank: window too large. Guard narrow crops before calling it. |
extract_area |
The degenerate-box guard compared the trimmed ink box against its own crop width. A crop is exactly the glyph’s columns, so every real glyph filled its own crop and every glyph was discarded. Compare against the frame. |
join |
join(in, direction, extend: 255) is a syntax error at runtime (unknown option extend); the rubyvips signature is expand: plus background:. |
The first two are worth remembering beyond captchas: any time you hand find_trim a mask you built yourself, pass background:.
What the numbers actually say
eval.rb generates 300 captchas, solves each one, and compares against the answer it generated. It also runs the same 300 through a baseline: a plain global cut, anything darker than the cut is ink, no reference image.
=== background subtraction, threshold 30 ===
captchas 300
submitted 298 (99.3%)
abstained 0 wrong segment count, 2 low confidence
character accuracy 100.0% of submitted
full-string accuracy 100.0% of submitted
solving 300 captchas took 10.57s (35.2ms each)
=== background subtraction, templates from clean renders ===
captchas 300
submitted 297 (99.0%)
character accuracy 100.0% of submitted
full-string accuracy 100.0% of submitted
=== global cut=130 ===
captchas 300
submitted 2 (0.7%)
abstained 181 wrong segment count, 117 low confidence
full-string accuracy 0.0% of submitted
=== global cut=170 ===
submitted 0 (0.0%)
abstained 285 wrong segment count, 15 low confidence
nothing was submitted
Three honest readings:
The baseline is not “worse”. It is dead, and the reason is the generator. Ink spans 30–170 levels of contrast while paper spans 84 levels end to end, so ink and paper overlap in absolute grey. No cut exists that separates them. If your captchas are a single flat tone with solid black digits, background subtraction buys you nothing. Mine is the hard case on purpose.
100% is an upper bound, not a field measurement. Templates come from the same generator as the evaluation set, the font is a crisp bitmap, and the tilt is ±6°. The middle run is the sanity check that matters: templates built from clean upright renders, one per digit with no noise and no tilt, still score 100%. So the tolerance comes from averaging and normalisation, not from leaking tilt into the templates. I would not expect that to survive anti-aliased glyphs at a different resolution.
Abstaining is the feature, not the failure. Two of 300 captchas came back low-confidence and were thrown away rather than submitted. A wrong guess costs a real attempt on the login endpoint:
def solve(image, solver: Solver.method(:new))
glyphs = solver.call(image).normalized
return [ :abstain_segments, "" ] unless glyphs.size == Sample::LENGTH
text = +""
glyphs.each do |glyph|
char, _distance, margin = CLASSIFIER.match(glyph)
return [ :abstain_margin, text ] if margin < MIN_MARGIN
text << char
end
[ :success, text ]
end
Two ways to abstain: wrong segment count (digits fused or a stroke vanished) and low margin (two templates too close to call). Either way, fetch a fresh captcha. Measured over 60 logins that allow 5 fetches each:
| Arm | Connected | Captchas fetched per login |
|---|---|---|
| Background subtraction | 60/60 (100%) | 1.00 |
| Same, clean templates | 60/60 (100%) | 1.02 |
| Global cut=130 | 0/60 (0%) | 5.00 |
At ~35ms per captcha, five attempts cost 175ms of CPU. No reason to gamble on a guess when the server hands you a new challenge for free.
The production solver looks nothing like this
The captcha I actually have to beat is not six stamped digits. It is 34 characters (ten digits and 24 lowercase letters, no l and no z), rendered with anti-aliasing, always at the same place inside the same crop window (CROP_X 75, CROP_Y 24, CROP_WIDTH 130, CROP_HEIGHT 35), over a background image I have vendored as empty.jpg. Pixel-exact templates do not survive that.
The solver I ship keeps stages 1 and 2 (subtract, amplify the difference ×10, threshold, split on blank columns) and throws away stage 4. Instead of comparing pictures, it measures five numbers per glyph:
# app/models/captcha_solver.rb (excerpt; image/left/top/columns are instance state)
def features(glyph, height = CROP_HEIGHT)
[
image.sum { |row| row.sum } / 256.0, # total ink
image.sum { |row| row[0...left].sum } / 256.0, # ink left of centre
image[0...top].sum { |row| row.sum } / 256.0, # ink above centre
image[top...height].sum { |row| row.sum } / 256.0, # ink below centre
columns.to_f # glyph width
]
end
Five ink statistics, matched against hand-labelled reference tables with weights [1, 3, 2, 8, 3], minimum margin 12, maximum distance 60. Deliberately coarse: total ink, where the mass sits, how wide the glyph is. Anti-aliasing and JPEG ringing move the pixels; they barely move the mass.
The retry loop:
# app/models/account.rb
while attempts < max_attempts
attempts += 1
captcha = client.captcha # a fresh challenge every time
result = solve_captcha(solver, captcha.png_bytes)
next unless result&.ok? # abstaining costs nothing
begin
session = client.authenticate(captcha_id: captcha.id, captcha_text: result.value)
return finish_connect!(client, session, via: "solver")
rescue Client::RequestFailed => error
raise unless error.code == Client::WRONG_CAPTCHA_CODE
end
end
raise Client::NeedsHuman, "Automatic captcha solving gave up after #{attempts} attempts."
Submit only confident solves, retry only on a wrong-captcha response, fall back to the human screen instead of hammering the endpoint. auto_create in the controller is that last line in one route.
For the record: this is for logging into an account I own, from a server I control. A captcha is an anti-automation control on someone else’s property. The interesting part of this post is the image arithmetic, not the endpoint.
Takeaway
Four libvips calls (subtract/abs, relational_const, reduce, resize) plus an averaged template and a margin turn a noisy 290x80 PNG into a string in tens of milliseconds, with no ML and no network.
The two ideas which increased the accuracy were not just trusting the pixels:
-
subtract the background so the lighting cancels, average ~48 renders into one soft template per digit, and refuse to answer when the runner-up is too close.
-
When the real captcha refused pixel templates, the fix was not a better algorithm. It was asking for less: five numbers about where the ink is, instead of a picture of it.
Everything here is runnable: sample.rb, pipeline.rb, eval.rb, figures.rb, MIT licensed at github.com/zoras/create-and-solve-captcha-with-rubyvips
. Every number above comes from a seed; ruby eval.rb reproduces the whole table.