The GDELT Project

Comparing Gemini 3, 3.1, 3.5, 3.6 Pro, Flash & Flash Lite For Creating a Likelihood-Impact-Triggers-Mitigations Analysis Of Iranian News

What would it look like to run a deep Likelihood-Impact-Triggers-Mitigations analysis over a day of Iranian television news on PressTV from July 13th using Gemini 3.1 Pro ($0.22), Gemini 3 Flash ($0.04), Gemini 3.1 Flash Lite Minimal Thinking ($0.04), Gemini 3.1 Flash Lite High Thinking ($0.04), Gemini 3.5 Flash ($0.14), Gemini 3.5 Flash Lite Minimal Thinking ($0.02), Gemini 3.5 Flash Lite High Thinking ($0.03), Gemini 3.6 Flash Medium Thinking ($0.14) and Gemini 3.6 Flash High Thinking ($0.14)? The numbers in parentheses are the cost of each analysis, combining input, thinking and output tokens. No data was used to train or tune any model.

Let's use the following prompt for all of the analyses below.

Read through the attached television news transcripts.
Catalog all of the potential geopolitical, economic, health, etc risks mentioned or suggested by the text.
Organize them as a likelihood + impact matrix that rates each for their likelihood of happening and the impact of that event if it does happen, along with likely triggers and recommended mitigations.
Output this as a bulleted list with the risk in bold and a concise summary of the likelihood, impact, potential triggers and recommended mitigations for each as text in the bullet, with each bolded.

Gemini 3.1 Pro ($0.22)

Gemini 3 Flash ($0.04)

Gemini 3.1 Flash Lite Minimal Thinking ($0.04)

Gemini 3.1 Flash Lite High Thinking ($0.04)

Gemini 3.5 Flash ($0.14)

Gemini 3.5 Flash Lite Minimal Thinking ($0.02)

Gemini 3.5 Flash Lite High Thinking ($0.03)

Gemini 3.6 Flash Medium Thinking ($0.14)

Gemini 3.6 Flash High Thinking ($0.14)

Finally, let's ask Gemini 3.1 Pro to compare all of the results:

The same input text and prompt were run through various Gemini models.
The outputs are below.
Compare the outputs and evaluate how the models compare against each other for this analysis.

The evaluation reveals distinct differences in information extraction, tone adherence, structural formatting, and the impact of "Thinking" levels.

Here is a detailed comparative analysis of the models:

1. Information Extraction & Granularity (The "Detail Test")

To test how well the models read the text, we can look at specific "easter eggs" (highly specific data points or names) embedded in your original source text and see which models caught them:

Winner for Granularity: Gemini 3.6 Flash (High Thinking) and Gemini 3.1 Pro.

2. Tone and Contextual Fidelity

The source text clearly contains specific ideological phrasing (e.g., "mini-genocide," "Zionist aggression," "scholasticide").

Winner for Tone/Context: Gemini 3 Flash. It acted most like a true analyst reading between the lines of the provided text.

3. The Impact of "Thinking" Levels (Minimal vs. Medium vs. High)

The prompt provides a great case study on how adding "Thinking" changes the output of the budget (Lite) and newer (3.6) models.

4. Formatting and Readability

Summary Verdict & Cost-to-Performance Ratio

  1. Best Overall Output (Quality + Formatting): Gemini 3.6 Flash High Thinking ($0.14). It found details every other model missed (the 50°C heatwave, 20% oil stat, Ras Tanura, the Paz administration) and formatted it perfectly. It is well worth the $0.14 price point for complex intelligence analysis.
  2. Best Value / Highest ROI: Gemini 3 Flash ($0.04). At a fraction of the cost of the Pro/3.6 models, it provided incredibly deep, context-aware analysis, complete with parenthetical justifications and exact thematic extraction ("scholasticide").
  3. The "Smart but Sanitized" Award: Gemini 3.1 Pro ($0.22). It was the only model brave enough to output "Trump," showing it relies less on safety-filter sanitization for political figures, but it is the most expensive and missed the climate/heatwave details.
  4. Models to Avoid for Deep Analysis: Gemini 3.1 and 3.5 Flash Lite (Minimal Thinking). At 0.02–0.04, they are cheap, but they generalize the data so much that the unique intelligence value of the source text is completely lost. If you must use Lite, always use High Thinking.