- AI Fire
- Posts
- ⭐ Google Just Released 3 New AI Models, But The One Everyone's Waiting For is Still MIA
⭐ Google Just Released 3 New AI Models, But The One Everyone's Waiting For is Still MIA
Learn which model to try now, how to use it effectively, and get a head start on workflows before the main model is released.

TL;DR
Google released 3 new Gemini models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber.
The best model to start with today is Gemini 3.6 Flash. It gives the strongest balance of quality, speed, and cost.
Gemini 3.5 Flash-Lite is better for cheap, high-volume work. Gemini 3.5 Flash Cyber is the security-focused model, but it is restricted to governments and trusted partners.
Key points
Gemini 3.6 Flash uses around 97 tokens vs. 276 tokens for Gemini 3.5 Flash on deep software engineering.
Avoid picking Flash-Lite only because it is cheaper.
Use Flash-Lite only when volume and token cost matter more than answer quality.
Table of Contents
What matters most when you choose an AI model? |
Introduction
So, Google DeepMind just released 3 new Gemini models.
And no, this still isn't Gemini 3.5 Pro. I know a lot of you guys were waiting for Google's next big Pro model. But instead, Google gave us a Flash-series refresh.
So here's the hands-on comparison I'll walk you through:
Gemini 3.6 Flash vs. Gemini 3.5 Flash-Lite
Speed, price, coding, long-context work, agent workflows, and real-world use cases
By the end, you'll know which of the new Gemini models actually fits your workflow, which one gives better value, and why this launch feels more like Google buying time before Gemini 3.5 Pro and Gemini 4.
I. Gemini 3.6 Flash: The Workhorse
Gemini 3.6 Flash is the main model in this release, and honestly, it's the one that gets the most practical upgrades.
It's the stronger upgrade from Gemini 3.5 Flash, especially for coding, knowledge work, multimodal tasks, and agent workflows.
Gemini 3.6 Flash is built for people who want better quality without moving to a slower, more expensive Pro-style model.
1. What Gemini 3.6 Flash Is Built For
Gemini 3.6 Flash sits right in the middle of Google's lineup. It's not the cheapest model, that job belongs to Gemini 3.5 Flash-Lite. But it's also not the heavy Pro model a lot of people are still waiting on.
Its job is to give you a better balance between:
Coding quality, document understanding, financial analysis, and multimodal work
Speed, output-token efficiency, and cheaper cost per agent task
In agent workflows, the model reads files, calls tools, checks results, retries, writes code, edits the plan, and runs again. So a small token saving per step can turn into a real cost difference after 20 or 50 steps.
2. Big Upgrade: Better Results with Fewer Tokens
Gemini 3.6 Flash doesn't just score higher, it also uses fewer output tokens than Gemini 3.5 Flash to get there.
Area | Gemini 3.6 Flash result | Compared with Gemini 3.5 Flash |
|---|---|---|
Output token efficiency | Lower cost per completed task | |
DeepSWE token reduction | Much better for coding agents | |
Input price | $1.50 / 1M tokens | Same input price as 3.5 Flash |
Output price | Cheaper than 3.5 Flash output |
This is exactly why Google is pushing it as an agentic model.
If your AI agent needs 10 rounds to finish a task, Gemini 3.6 Flash should spend fewer tokens while giving stronger answers than Gemini 3.5 Flash.
3. Benchmark Results: Where 3.6 Flash Improves
Here's the clean comparison against Gemini 3.5 Flash:
Benchmark | Gemini 3.6 Flash | Gemini 3.5 Flash | What it means |
|---|---|---|---|
DeepSWE | 49% | 37% | Better at software engineering tasks |
MLE Bench | 63.9% | 49.7% | Better at machine learning engineering |
OSWorld-Verified | 83.0% | 78.4% | Better at computer-use tasks |
GDPval-AA v2 | 1421 | 1349 | Better at knowledge work |
For normal users, Gemini 3.6 Flash should be better when the task needs reasoning across messy information.
That could mean reading a long financial report, moving code from one framework to another, finding problems in a repo, or extracting useful details from charts, PDFs, screenshots, and transcripts.
4. Where I Would Actually Use Gemini 3.6 Flash
I'd use Gemini 3.6 Flash as the main "do the hard work" model in a Gemini workflow. Some practical examples:
Use case | Why 3.6 Flash fits |
|---|---|
Financial data analysis | It can read tables, reports, transcripts, and summarize key changes |
Code migration | Better DeepSWE and MLE Bench scores make it stronger for engineering work |
Multimodal document parsing | Useful when your input has text, charts, screenshots, or mixed file types |
Internal research agents | Stronger knowledge-work score helps with multi-step analysis |
Product and design workflows | Customers like Figma, Harvey, Hebbia, and JetBrains are already cited around this release |
Here's a prompt I'd actually use with it:
I’m going to give you a long business document with tables, charts, and notes.
Your job is to analyze it like a senior analyst. First, summarize the 5 most important findings. Then explain what changed, what risks I should pay attention to, and what follow-up questions I should ask.
Keep the answer practical. If a number is important, mention where it appears and why it matters.
This document is 285 pages long, and it takes about 30 seconds to generate a result.
For reference: Gemini 3.6 Flash is the safer pick when the task has enough complexity that a weaker model might need extra retries to get it right.
5. Safety and Availability
Google also added stronger Frontier Safety safeguards to Gemini 3.6 Flash.
That means the model is designed to reduce risky outputs in areas such as CBRN and cyber-offense misuse (CBRN means chemical, biological, radiological, and nuclear risk areas.)
Cyber-offense misuse means prompts that try to push the model toward harmful hacking behavior.
It's live now through: Gemini API via AI Studio, Android Studio, Google Antigravity, Gemini Enterprise, and the Gemini app.

6. My Take on Gemini 3.6 Flash
Gemini 3.6 Flash is the safest default choice from this release.
It gives you stronger coding results, better knowledge work, better computer-use performance, and lower output cost than Gemini 3.5 Flash.
So use Gemini 3.6 Flash when quality still matters, but you also care about speed and agent cost.
Learn How to Make AI Work For You!
Transform your AI skills with the AI Fire Academy Premium Plan - FREE for 14 days! Gain instant access to 700+ AI workflows, advanced tutorials, exclusive case studies and unbeatable discounts. No risks, cancel anytime.
II. Gemini 3.5 Flash-Lite: The Speed & Cost Option
Gemini 3.5 Flash-Lite is the budget model in this batch of new Gemini models.
This is the one you choose when you care more about throughput, cost-per-token, and volume than getting the strongest answer every time.
→ Gemini 3.5 Flash-Lite is built for fast, repeated work at scale.
1. What Gemini 3.5 Flash-Lite Is Built For
Gemini 3.5 Flash-Lite is Google’s fastest and most cost-effective 3.5-class model. It is made for workflows where the model needs to run again and again:
agentic search, document processing, extraction, classification, and quick computer-use steps
smaller background tasks inside larger agent workflows
That makes it useful for teams building AI agents, internal tools, research systems, customer support workflows, and document pipelines.
The key number here is speed. According to Artificial Analysis, Gemini 3.5 Flash-Lite can reach 463.7 output tokens per second.
That speed matters when you run hundreds or thousands of small tasks. A slower model may feel fine in one chat, but in a production workflow, latency stacks up fast.
2. Pricing: Why Flash-Lite Is So Cheap to Run
Gemini 3.5 Flash-Lite is much cheaper than Gemini 3.6 Flash.
Model | Input price | Output price | Best use |
|---|---|---|---|
Gemini 3.5 Flash-Lite | $2.50 / 1M tokens | High-volume tasks | |
Gemini 3.6 Flash | $1.50 / 1M tokens | $7.50 / 1M tokens | Better reasoning and harder tasks |
This is the main reason Flash-Lite exists.
If your workflow sends a lot of small prompts, checks many documents, or runs many agents at the same time, the price gap becomes very real.
→ Use Flash-Lite when the task is simple enough that paying for a stronger model would waste money.
3. Benchmark Results: Better than the Old Lite Model
Flash-Lite is cheaper, but it still improves a lot over Gemini 3.1 Flash-Lite.
Benchmark | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite | What it means |
|---|---|---|---|
Terminal-Bench 2.1 | 54% | 31% | Much better at terminal and coding-style tasks |
GDM-MRCR v2 | 72.2% | 60.1% | Better long-context handling |
GDPval-AA v2 | 1140 | 642 | Stronger at knowledge work |
The jump on GDPval-AA v2 is the one I care about most for business users.
It means Flash-Lite should be much better at practical knowledge work than the older Lite model. That includes reading internal docs, checking reports, pulling useful facts from long files, and helping with simple research flows.
4. Surprising Part: It Beats Gemini 3 Flash in Some Areas
Even though it is the cheaper model, Gemini 3.5 Flash-Lite beats Gemini 3 Flash on two useful benchmarks:
Benchmark | Gemini 3.5 Flash-Lite | Gemini 3 Flash |
|---|---|---|
SWE-Bench Pro | 54.2% | 49.6% |
OSWorld-Verified | 74.0% | 65.1% |
That does not mean Flash-Lite is the best coding model here. Gemini 3.6 Flash is still the better pick for more serious engineering work.
But it does mean Flash-Lite is strong enough for a lot of smaller coding, terminal, and computer-use tasks where speed matters.
5. Adjustable Thinking Levels
One useful part of Gemini 3.5 Flash-Lite is that you can adjust its thinking level.
For simple, high-volume work, you can keep thinking at minimal or low. That helps keep the model fast and cheap.
For harder subagent tasks, you can raise the thinking level so Flash-Lite spends more effort on multi-step work.

This gives you more control over cost and quality.
A simple setup would look like this:
Task type | Thinking level |
|---|---|
Classification, extraction, quick search | Minimal or low |
Multi-step subagent work | Medium or higher |
III. Gemini 3.6 Flash vs. Gemini 3.5 Flash-Lite
Among the new Gemini models, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are the two most people can actually use. Flash Cyber is restricted, so for developers, teams, and normal builders, this is the comparison that matters.
-> The simple answer: use Gemini 3.6 Flash when the task needs judgment. Use Gemini 3.5 Flash-Lite when the task needs scale.
Choose this model | When your task looks like this |
|---|---|
Gemini 3.6 Flash | The task is complex, messy, expensive to get wrong, or needs stronger reasoning |
Gemini 3.5 Flash-Lite | The task is repetitive, high-volume, speed-sensitive, or cheap enough to run many times |
That is the easiest way to avoid overpaying.
The strongest setup is using both together. Here is how that can look in a real workflow:
Workflow step | Best model |
|---|---|
Understand the goal | Gemini 3.6 Flash |
Break the task into smaller jobs | Gemini 3.6 Flash |
Generate many drafts, searches, or extractions | Gemini 3.5 Flash-Lite |
Filter weak outputs | Gemini 3.5 Flash-Lite |
Review, refine, and make the final decision | Gemini 3.6 Flash |
This pairing makes sense because the expensive part should be the thinking.
IV. Gemini 3.5 Flash Cyber: The Restricted One
Gemini 3.5 Flash Cyber is the hardest model to judge in this release because most people cannot use it.
This is one of the new Gemini models, but it is not a normal public chat model. It is a specialized version of Gemini 3.5 Flash, fine-tuned for cybersecurity work.
-> Its job is simple: find and fix security vulnerabilities in code.
Area | What it means |
|---|---|
Base model | Fine-tuned from Gemini 3.5 Flash |
Main use | Vulnerability discovery and repair |
Deployment | Runs inside CodeMender |
Access | Governments and trusted partners only |
Public testing | Not available right now |
The important part is how Google deploys it.
Gemini 3.5 Flash Cyber runs inside CodeMender, Google’s AI code-security agent. Multiple Flash Cyber agents can work together, inspect code, search for weak points, and produce one vulnerability report.
That makes it very different from Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.
Those two models are built for broad usage. Flash Cyber is built for a narrow, high-risk job.
From a security point of view, that makes sense. A model that can find and fix vulnerabilities at a high level is useful, but it is also dual-use. In the wrong hands, the same skill can create real damage.
→ My take: Gemini 3.5 Flash Cyber is probably the most important model in this launch for security teams, but it is also the least accessible for normal developers today.
Conclusion
The new Gemini models are useful, but this still feels like a Flash-series refresh before Google’s bigger move.
Gemini 3.5 Pro is still not public, and Gemini 4 is already further ahead in Google’s pipeline.
So my final take is simple:
→ Use Gemini 3.6 Flash as your main model. Use Gemini 3.5 Flash-Lite when you need speed and scale. Keep Gemini 3.5 Flash Cyber on your watchlist, because it may matter a lot for security teams if Google expands access later.
If you are interested in other topics and how AI is transforming different aspects of our lives or even in making money using AI with more detailed, step-by-step guidance, you can find our other articles here:
ChatGPT Work for Marketers: Complete Guide to Turn ChatGPT Into a Real Workflow Agent
Sonnet 5 vs Opus 4.8: Same Result, 50% Cheaper? Read This Before Switching
4 Vibe Coding Use Cases That Let You Earn With ANY AI Model in 2026 (Fable, GPT Sol, Kimi)*
Kimi K3 IS INSANE! Best Open Model EVER That BEATS FABLE 5 & GPT-5.6! (Fully Tested)*
5 Claude AI Skills That Turn Your ALL Social Media Data Into Decisions*
*indicates a premium content, if any







Reply