My machine is a living life. I’ll prove it.
Introduction

hCaptcha has just received another update today, introducing a new tag that requires selecting in the sky left-flying airplanes.
Fortunately, however, every sample image contains an airplane, so this tag effectively removes one constraint, leaving only the two constraints of in the sky and flying left.
Main Method
First, let’s observe the images, still referring to the collected dataset1. Besides the fact mentioned earlier that each sample image contains an airplane, if the airplane is in the sky, its background must be very “clean”. If it is not in the sky, it can basically be judged as being on the ground. Images on the ground also consist of multiple regions, such as lawns, airport runways, background forests, background skies, etc.
So, the key to distinguishing the first problem is the complexity of the background. How to do it?
My initial idea was to inherit the previous approach of color block filtering, but after dividing the color blocks, it was difficult to distinguish whether a block was an airplane or a background block, so it was discarded.
Then, I looked at most methods for removing sky backgrounds, which are basically based on threshold filtering in the HSV color space, followed by morphological operations like erosion and dilation for noise reduction. I tested it a few times myself, but the color range of the sky is slightly too large, containing blue, white, and yellow. More critically, it is very similar to the color of airplanes, as airplanes are mostly light colors like pale blue or white. This was also discarded.
I then thought that the shadow under the airplane would be black, so creating a superpixel smart selection based on black could work, but I don’t know how to implement it (x, so it was discarded.
Finally, the adopted solution was to use the contour line method Canny to find all contour lines, then set a threshold to determine if the airplane is in the sky based on the number of contour lines. If it is in the sky, the contour lines will be very simple, whereas if it is not in the sky, a large amount of chaotic lines will be added. Of course, this threshold was derived by randomly testing a few images, given the huge difference between the two cases. After such processing, the judgment accuracy approaches 100%.
Okay, one problem solved. Now, the remaining problem is: how to determine if the airplane is facing left?
This really stumped me. Without using Deep Learning, it is indeed difficult, but there are some clever tricks. First, most airplanes are transport planes, fighters, or passenger jets; there are rarely propeller planes in the front, and I haven’t seen helicopters either. There are even WTF Airplanes like the one below (what the heck is this?)

Since that’s the case, these types of airplanes have a characteristic: “light head, heavy tail”. Besides the tail fin being heavier and the nose being pointed and lighter, the wings also point backward, so the “center of gravity” of the entire image should be biased towards the tail. I could determine the airplane’s direction based on 4 points: extreme left, extreme right, midpoint, and center of gravity.
Sounds scientific, right? However, the actual effect was not very good. The most critical issue is that airplanes have perspective relationships, so from the front view, the center of gravity might appear to be at the front. For example, if the nose is facing you, a large number of lines are drawn above the nose, while there are very few lines at the tail, making the center of gravity of the entire image appear at the front. Later, I wondered if I could fill the image, but the resulting contour lines were mostly not closed, making it difficult to fill the entire airplane (if that were possible, image segmentation would be simple).
Later, I had a sudden inspiration and took a different path. Still observing the contour lines drawn above, the nose, lacking complex elements, produces relatively “simple” contour lines, while the tail, due to components like “tail fins”, produces relatively “complex” ones. So how to measure this “simplicity” and “complexity”? Just sum them up… That’s right, in the end, I counted the non-zero pixels from x_min to x_min + left_threshold on the left (which are the contour line pixels) and compared them with the pixel count from x_max - left_threshold to x_max. Whichever is larger is the tail; if the right side is larger, then the nose is on the left. Unexpectedly, the final result was quite good, basically passing verification within 1-2 rounds.
Attached is the complete code for the test version
| |
Conclusion
To be honest, solving a high-quality image processing problem from hCaptcha every day still feels pretty cool hhhh.
However, through sharing by netizens, I have seen more problems solved using generative models. After all, they are paid to do this, and later some tasks can no longer be solved just by image processing, such as black-and-white striped cats that appeared in feedback for certain Tampermonkey plugins.
Just a clever workaround to meet each challenge.

