What Those Few Seconds of "Select the Traffic Lights" Leave Behind
"Select all the traffic lights." "Click on images containing crosswalks." - This image verification that appears every time you log into a website. Annoying, right? But beyond "proving you're not a robot," those tasks leave something behind.
Your clicks accumulate as labels that only a human can supply: "this image contains a traffic light." What those labels are actually used for is where verifiable fact and internet folklore get tangled together, so that is where this article starts.
The Evolution of CAPTCHA
CAPTCHA emerged in the early 2000s as a "test to distinguish humans from computers." Starting with distorted text recognition, it has evolved alongside advancing technology.
- 1st generation (2000s): Type distorted characters. Became breakable as OCR technology improved
- 2nd generation (reCAPTCHA, acquired by Google in September 2009): Had humans read words from scanned old books that OCR couldn't decipher. Google titled its acquisition announcement "Teaching computers to read" and stated outright that the answers would go toward digitizing books
- 3rd generation (reCAPTCHA v2, 2014-): "I'm not a robot" checkbox + image selection. Traffic lights, crosswalks, buses, and bicycles appeared
- 4th generation (reCAPTCHA v3, 2018-): Scores user behavior patterns and only shows image verification for suspicious cases. Most users pass without seeing a CAPTCHA
Why "Traffic Lights" and "Crosswalks"?
What reCAPTCHA v2 puts in front of you is almost always something you would find in a street scene: traffic lights, crosswalks, buses, bicycles, fire hydrants. People noticed that this overlaps with the objects a self-driving car has to recognize, and from there the claim "those tasks are building training data for autonomous vehicles" spread widely.
No primary source backs that claim up. Google acquired reCAPTCHA in September 2009, and Waymo was established in 2016 as a self-driving technology company under Alphabet, so the timelines do overlap. But Waymo has never named reCAPTCHA as a source of its training data, and Google has never said that CAPTCHA answers feed autonomous driving. Between "street scenes get used as challenges" and "this is autonomous driving training data" sits a leap that nobody has actually filled in.
What is certain is that your answer persists as a human-supplied label, and that this work happens unpaid, all over the world, every day. What it gets used for has not been published.
What's Behind "I'm Not a Robot"
Sometimes you can pass by simply clicking the "I'm not a robot" checkbox. So what is being looked at in that brief moment? This is the part that gets misdescribed most often.
Google's own documentation goes no further than saying that the score is based on interactions with your site and that reCAPTCHA learns by seeing real traffic on that site. Which signals are weighted, and how, is not published. Explanations along the lines of "it watches how your mouse wobbles" or "it measures how far off-center you clicked" circulate widely, but no primary source from Google says so. Publishing the internals would simply tell bot authors what to imitate.
Some things do follow from the shape of the mechanism. reCAPTCHA runs as a Google script embedded in the page, so its requests carry your IP address, exactly as any other web request does. Information sent by your browser is unavoidably part of the picture, whether or not the weighting is disclosed. And cookies remain the most basic identification mechanism on the web.
If the verdict is "human-like," the checkbox alone is sufficient; if it is "suspicious," image selection appears. Being shown images does not mean you did anything wrong. One of the primary threats CAPTCHA defends against is credential stuffing, where bots use stolen passwords to attempt mass logins.
The Future of CAPTCHA - Invisible Authentication
reCAPTCHA v3, released in October 2018, returns a score in the background without requiring any user interaction. Here too, the official description stops at "the score is based on interactions with your site"; what it actually looks at is not disclosed. Deciding what to do with that score - let the request through, ask for extra confirmation, or refuse it - is left to the site that installed it.
In the future, CAPTCHAs may become completely "invisible," with authentication completing without users even being aware of it. For developers building web services, understanding API security fundamentals is essential to implementing bot protection on the server side as well.
Summary
A CAPTCHA image challenge is a machine for producing, free of charge, labels that only humans can supply. But the familiar explanation that "those are autonomous driving training data" has no primary source behind it, and neither Google nor Waymo has ever said it. What is being looked at behind "I'm not a robot" is undisclosed as well; what is certain is that your IP address reaches the other side, exactly as it does in any other web request. You can see how that IP address looks from outside at IP確認さん. Next time a CAPTCHA appears, consider not what those few seconds are for, but what they definitely leave behind.
Related Terms in This Article
Frequently Asked Questions
Why does CAPTCHA ask me to select traffic lights and crosswalks?
Beyond verifying you are human, the task collects image labels that only a human can supply. Because traffic lights and crosswalks are what you get asked about, the story that this builds self-driving training data spread widely, but no primary source supports it and neither Google nor Waymo has ever said so officially.
Why can I sometimes pass with just the checkbox?
Google's documentation goes no further than saying the score is based on interactions with your site, so which signals are weighted, and how, is not published. If the verdict is "human-like," the checkbox alone is enough; if it is "suspicious," image selection is added.
What is reCAPTCHA v3?
reCAPTCHA v3 runs entirely in the background without any user interaction and returns a score, where 1.0 is very likely a good interaction and 0.0 is very likely a bot. Which signals produce that score is not published. The site that installed it decides what to do with the result.