Site icon The GDELT Project

Experiments With Detecting & Correcting AI Hallucination In Geographic Mapping: Testing Gemini 3.8 Flash

Earlier today we explored how Gemini 3.7 Flash was able to readily identify key hallucinations in an AI-generated map, though it wasn't perfect. Given the enormous leaps in visual reasoning we've seen from its successor Gemini 3.8 Flash, does it majorly improve on Gemini 3.7's performance? On this task, Gemini 3.8 Flash seems to perform on par with 3.7, struggling in particular to follow the lines from each label back to the source country. For example, it correctly assesses that Aruba (bottom left) is in the wrong place, but assesses that it points to Chile or Peru, not Argentina. Strangely, it flags that China points to the country country, because it believes the line terminates in Russia and can't follow it all the way to where it accurately points to China. The results here both suggest that visual reasoning models can at least partially flag errors in AI-produced maps, but also points to continued struggles in their ability to follow connective lines to fully understand where a given label points, which we've observed in past experiments.

Here is the prompt we used with the map above and the results:

Look closely at all of the labels in this map and make an EXHAUSTIVE list of ALL geographic errors.
Output as a bulleted list.

 

Exit mobile version