
Introduction
Recently, hCaptcha1 underwent an update, shifting from previous object recognition to selecting vertical rivers within images. As shown in the image above, screenshot from 1.
To be honest, the first glance revealed it was done by some generative models because the image features are obvious: some areas are particularly blurry, and the generation logic is quite nonsensical. Initially, I thought it was an application of GPT-3, converting image descriptions into text, but later deemed it unreasonable due to the lack of imagination. Moreover, the distribution of elements is extremely obvious; everything is straight and unnatural. It must be some work on generating raw images from semantic images, such as NVIDIA’s artistic creations.
Before starting, I never imagined that NVIDIA’s GANs could be applied to the CAPTCHA field, nor did I expect that after the upgrade, it would be downgraded from object detection to requiring only image processing to pass.
Main Content
@QIN2DIM previously wrote about a hCaptcha-challenger2 project using YOLOv5 to handle this. After the update, they released a dataset vertical_river3, so I used image processing to quickly implement a Demo.
First, the image features are very obvious, mainly consisting of several elements: sky, mountains, grass, and water in the background, with very distinct divisions and distributions, all appearing as separate blocks.
My initial idea was to apply some filters, such as 高斯滤波, and then directly use slic for superpixel segmentation. The expected result was that superpixels would group each color block into a single superpixel. However, the final result was not ideal because superpixel segmentation relies heavily on factors other than color, and limiting the number of superpixels causes larger color ranges to be grouped into one class. Here is an example.

As you can see, there are many issues: some parts that should be segmented were not correctly divided and instead clustered together, while some parts that shouldn’t be segmented were clustered together due to the filtering effect.
After repeatedly adjusting parameters with no success, I felt that this segmentation approach should not treat a single superpixel as a color block, but rather use the superpixel segmentation results as a reference or boundary for segmentation.
I originally wanted to try some edge enhancement algorithms but couldn’t find a suitable one because I observed that most CAPTCHA images have very blurry boundaries. Therefore, preserving edge information during segmentation is crucial. When selecting filters, it is essential to prioritize whether the filter can correctly preserve edge information rather than blurring the edges. Among them, 双边滤波 is an excellent choice, and 均值偏移 is also a good option.
After these two steps, you have become myopic, but the edges remain relatively clear. Now you can proceed with color block segmentation. Here, I mainly used scikit-image.graph.rag_mean_color to obtain average color blocks, primarily referencing the official RAG Merge4 implementation.
The judgment condition was written simply: I only checked if the last row contains three or more color blocks. If so, I consider that the middle is separated by a river, indicating the presence of a vertical river. I feel this judgment condition could still be optimized, but after testing, out of about 100 images, only around 2 were misclassified. Even without optimization, this accuracy is acceptable.
The final effect is as follows:
I submitted a PR, and the effect after @QIN2DIM merged it is as follows:
Code used for testing and visualization:
| |
Conclusion
Neural networks have made everything more complex, yet also simpler.
(Discussing the experience of a hCaptcha engineer seeing their project, developed over half a day, surpassed by a much simpler algorithm in less than a night?)





