GPT-5.6 Sol closes OpenAI's vision gap, but Gemini stays cheaper and better
OpenAI’s new GPT-5.6 family — Sol, Terra, and Luna — marks the company’s first serious leap in visual understanding, according to benchmarks run by Roboflow ahead of its forthcoming VLM test suite. The gains are sharpest in object detection, where the flagship Sol jumped to 46.2 mAP@50 from GPT-5.5’s dismal 13.8, and in counting, which climbed from 64.9% to 73%. Even Luna, the cheapest model in the lineup, beats the previous generation. That matters because the release stream leaned heavily on computer-use agents, UI navigation, and 3D visualization — all of which hinge on the model actually seeing what’s on screen.
The improvements are uneven and come with caveats. OCR was essentially flat and targeted text extraction actually regressed (Sol scored 82.5% versus GPT-5.5’s 87.6%). OpenAI confirmed that Sol grows unstable on images around 2,000×2,000 pixels or larger, sometimes emitting bounding boxes in unnatural rows unrelated to the actual objects; higher reasoning effort fixes this but drives up cost and latency, so resizing or cropping is the practical workaround. Coordinate format is also finicky — GPT-5.6 wants absolute XYXY pixels, and using the wrong convention cost roughly 15 mAP points.
The headline caveat is economics. Sol averages ~10 seconds and ~2.5 cents per image, making it the second most expensive model tested after Claude Fable 5. Gemini 3.5 Flash still leads Roboflow’s detection and counting benchmarks at 0.8 cents per image, keeping it the stronger pick for high-volume workloads. The takeaway: OpenAI has moved vision from a weakness to a genuine capability and is now competitive for agents, document workflows, and screen understanding — but it hasn’t taken the price-performance crown.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.