Earlier today we demonstrated that Gemini's visual reasoning capabilities have advanced so far with its latest Gemini 3.8 Flash model that it appears we could now reliably detect broadcasts with corrupted or unusable video content simply by having Gemini visually examine a 1/4 fps thumbnail grid of the broadcast. In those examples we asked Gemini only to classify the video as a whole as corrupted and whether it contained valid programming. What if we ask Gemini something far harder: visually describe the entire broadcast from just this 1/4 fps thumbnail grid, row by row with sufficient visual and semantic detail that a researcher could visually search the broadcast for relevant snippets through a simply keyword search of these descriptions? The end result, as seen below, is extremely accurate, despite the coverage being from Afghanistan, which is likely poorly represented in the training datasets of most models and cost just just less than two pennies ($0.017) for an entire half-hour broadcast! No data was used to train or tune any model and only the 1/4 thumbnail grid was used for the analyses below.
What about asking Gemini to visually assess a bit about what the broadcast is showing as a whole?
This is a thumbnail grid of a video. Describe in a paragraph what this broadcast appears to cover and give it a title.
Amazingly Gemini describes it quite well:
Title: Maseer Panel Discussion on RTA: State and Current Affairs in Afghanistan
This broadcast from Radio Television Afghanistan (RTA) opens with a news montage highlighting national security forces, infrastructure projects, and government proceedings before transitioning into the current affairs talk show Maseer (The Path). Hosted in a modern studio, the program convenes a panel of Afghan officials, religious scholars, and political analysts dressed in traditional attire to debate and discuss pressing governance, social, and policy issues facing the Islamic Emirate of Afghanistan. Throughout the episode, the moderator directs discussion points to various seated panelists, allowing each guest dedicated time to present their perspectives on national developments.
How about a far more ambitious and complex task? Visually describing the entire broadcast in 24 second chunks by describing the 1/4fps thumbnail grid row by row? This would enable a form of "visual search" in which a researcher could keyword search this textual description to find segments of the broadcast whose textual visual description matches their keywords:
This is a thumbnail grid of a video. Describe row by row what you see and what the broadcast appears to be covering in that row. Output as a numbered list, with item 1 corresponding to row 1 and so on. Describe concisely but with enough detail that someone who can't see the thumbnail grid understands what it is showing and that would allow for visual search of the broadcast through keyword search of your descriptions. Describe not just what is visually being shown, but what is happening.
Amazingly, this rich visual description is quite accurate and cost just $0.017, suggesting this could be quite scalable:
- Row 1: News intro sequence depicting Afghan military vehicles on patrol, armed security personnel standing guard, and the blue national broadcaster (RTA/Melli) channel identity branding.
- Row 2: Channel logos, digital news clock, scenic Afghan desert landscapes, historical citadels, bustling bazaars, and national monuments.
- Row 3: Broadcast globe graphic, "Khabar Kootah" (News in Brief) title card, aerial cityscape, the presidential palace/Arg in Kabul, an iconic mosque, and armored vehicles navigating city streets.
- Row 4: Modern administrative building exteriors, city road traffic, pedestrian street crossings, news bumper graphics, and an indoor diplomatic meeting room.
- Row 5: High-level official meetings with Afghan elders and officials in traditional turbans, followed by press conferences at official government podiums.
- Row 6: A large indoor assembly/conference hall filled with seated Taliban officials, delegates, and dignitaries wearing turbans.
- Row 7: Military formation and inspection of uniform-clad and armed security personnel by Taliban commanders and officials.
- Row 8: Close-up portraits of prominent Taliban officials and religious leaders, transitioning to newly constructed multistory residential apartment complexes.
- Row 9: Crowds of local citizens gathered outdoors, view of modern housing developments, and urban traffic on city streets.
- Row 10: Public outdoor community gathering under an awning, displayed official government communique/decree document in Dari/Pashto, and daytime street scenes.
- Row 11: Everyday urban street views showing cars, taxis, motorcycles, auto-rickshaws, pedestrians, and traffic police managing traffic.
- Row 12: Exterior gate of an educational/administrative institution, institutional library interior with book displays, and an exhibition/reception room.
- Row 13: Red administrative school/complex buildings alongside modern industrial solar panel arrays and blue warehouse structures.
- Row 14: "Khabar Kootah" bumper transitioning into the RTA logo animation featuring glowing golden particle light bursts.
- Row 15: Glowing particle effects leading into the "Melli News" title sequence and large 3D blue Melli television logo.
- Row 16: Dynamic golden animated title sequence and bumper for the political talk show Maseer ("Path").
- Row 17: Talk show host wearing a dark vest and white traditional tunic seated at the studio desk, introducing the evening's discussion topic.
- Row 18: The host continues the program introduction before transitioning to introduce the studio guest panel.
- Row 19: Host speaks on camera, followed by wide panoramic studio shots displaying the host and panel seated around a round glass table.
- Row 20: Wide studio establishing shots cutting to a close-up of an elderly studio guest with a white turban and long white beard.
- Row 21: Close-ups of the elderly guest in white attire speaking and providing analysis with lower-third identification banners.
- Row 22: Elderly guest continues speaking, alternating with wide views of the circular studio stage.
- Row 23: The host interjects and poses questions, cutting between wide studio shots and close-ups of the elderly panelist.
- Row 24: Elderly guest explains his viewpoints, interspersed with medium and wide camera cuts of the panel discussion.
- Row 25: Sustained close-up shots of the elderly guest speaking emphatically on political/governance matters.
- Row 26: Wide angle of the studio set, the host addressing the panel, and the elderly guest gesturing while speaking.
- Row 27: Elderly guest continuing his commentary, with cutaways to the host taking notes and listening attentively.
- Row 28: Wide studio shots of the full discussion panel alternating with tight shots of the elderly speaker.
- Row 29: Elderly guest speaking, followed by the host preparing questions from his tablet and notes.
- Row 30: Host moderating and redirecting the discussion, wide stage views, and final remarks from the elderly guest.
- Row 31: Host turns the floor over to another guest, a bearded man wearing a black turban and dark vest.
- Row 32: Wide studio framing cutting to close-up shots of the guest in the black turban beginning his commentary.
- Row 33: Black-turbaned guest addressing the camera with lower-third title graphics, intercut with wide panel shots.
- Row 34: Close-ups of the black-turbaned analyst elaborating on political affairs, with cuts back to the host.
- Row 35: Host asking follow-up questions, cutting to the black-turbaned guest answering in detail.
- Row 36: Black-turbaned guest continuing his argument, host listening and moderating the debate.
- Row 37: Studio discussion featuring an information banner displayed along the bottom of the screen, focusing on the black-turbaned speaker.
- Row 38: Host interjecting, cutting to wide overhead and front studio angles of the three-person discussion.
- Row 39: Host posing another question to the black-turbaned guest, who continues discussing the topic.
- Row 40: Host transitions the conversation to a third studio guest, a man in a brown vest and white shirt.
- Row 41: Close-ups of the guest in the brown vest gesturing actively while articulating his perspective.
- Row 42: Close-up coverage of the guest in the brown vest speaking with hand gestures, with cutaways to the host.
- Row 43: Guest in the brown vest speaking, intercut with wide shots of the studio panel seated at the circular table.
- Row 44: Alternating shots between the host listening and the third guest in the brown vest making policy points.
- Row 45: Guest in the brown vest continuing his commentary, followed by host moderation and wide studio cuts.
- Row 46: Ongoing discussion by the guest in the brown vest, intercut with full-stage studio perspective.
- Row 47: Medium shots of the third guest speaking, cut to the host wrapping up this segment of questions.
- Row 48: Host directs questions back to the elderly white-bearded guest, with wide panoramic studio views.
- Row 49: Elderly guest speaking again with descriptive lower-third banners, intercut with the host listening.
- Row 50: Extended remarks from the elderly guest in white, alternating with the host moderating.
- Row 51: Studio wide views showing all panelists around the table, focusing on the elderly guest's closing points.
- Row 52: Close-ups of the elderly guest answering queries, followed by reaction shots from the host.
- Row 53: Wide views of the studio set as the elderly guest finishes his response and the host reviews his notes.
- Row 54: Host shifts the final round of questions to the black-turbaned analyst.
- Row 55: Black-turbaned guest delivering closing arguments, intercut with the host and wide studio setup.
- Row 56: Continued commentary by the black-turbaned panelist with camera cuts between medium and wide perspectives.
- Row 57: Host asking closing clarification questions as the black-turbaned guest responds.
- Row 58: Dialogue between the host and the black-turbaned guest, framed by full-studio table angles.
- Row 59: Black-turbaned guest providing final analysis with lower-third topic graphic displayed.
- Row 60: Host taking control of the show, wide studio panning shots, and final comments from the black-turbaned guest.
- Row 61: Concluding remarks from the black-turbaned analyst, with cuts to the host preparing to close.
- Row 62: Host concluding the debate and thanking the panel, with reaction shots from the black-turbaned guest.
- Row 63: Host delivering the program's outro, cutting between wide studio shots and panelists.
- Row 64: Final wide studio camera pull-back showing the entire set, host, and panel as the broadcast concludes.