• AI Fire
  • Posts
  • ⭐ Google Just Released 3 New AI Models, But The One Everyone's Waiting For is Still MIA

⭐ Google Just Released 3 New AI Models, But The One Everyone's Waiting For is Still MIA

Learn which model to try now, how to use it effectively, and get a head start on workflows before the main model is released.

TL;DR

Google released 3 new Gemini models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber.

The best model to start with today is Gemini 3.6 Flash. It gives the strongest balance of quality, speed, and cost.

Gemini 3.5 Flash-Lite is better for cheap, high-volume work. Gemini 3.5 Flash Cyber is the security-focused model, but it is restricted to governments and trusted partners.

Key points

  • Gemini 3.6 Flash uses around 97 tokens vs. 276 tokens for Gemini 3.5 Flash on deep software engineering.

  • Avoid picking Flash-Lite only because it is cheaper.

  • Use Flash-Lite only when volume and token cost matter more than answer quality.

What matters most when you choose an AI model?

Login or Subscribe to participate in polls.

Introduction

So, Google DeepMind just released 3 new Gemini models.

And no, this still isn't Gemini 3.5 Pro. I know a lot of you guys were waiting for Google's next big Pro model. But instead, Google gave us a Flash-series refresh.

So here's the hands-on comparison I'll walk you through:

  • Gemini 3.6 Flash vs. Gemini 3.5 Flash-Lite

  • Speed, price, coding, long-context work, agent workflows, and real-world use cases

By the end, you'll know which of the new Gemini models actually fits your workflow, which one gives better value, and why this launch feels more like Google buying time before Gemini 3.5 Pro and Gemini 4.

I. Gemini 3.6 Flash: The Workhorse

Gemini 3.6 Flash is the main model in this release, and honestly, it's the one that gets the most practical upgrades.

It's the stronger upgrade from Gemini 3.5 Flash, especially for coding, knowledge work, multimodal tasks, and agent workflows.

Gemini 3.6 Flash is built for people who want better quality without moving to a slower, more expensive Pro-style model.

1. What Gemini 3.6 Flash Is Built For

Gemini 3.6 Flash sits right in the middle of Google's lineup. It's not the cheapest model, that job belongs to Gemini 3.5 Flash-Lite. But it's also not the heavy Pro model a lot of people are still waiting on.

Its job is to give you a better balance between:

  • Coding quality, document understanding, financial analysis, and multimodal work

  • Speed, output-token efficiency, and cheaper cost per agent task

In agent workflows, the model reads files, calls tools, checks results, retries, writes code, edits the plan, and runs again. So a small token saving per step can turn into a real cost difference after 20 or 50 steps.

2. Big Upgrade: Better Results with Fewer Tokens

Gemini 3.6 Flash doesn't just score higher, it also uses fewer output tokens than Gemini 3.5 Flash to get there.

Area

Gemini 3.6 Flash result

Compared with Gemini 3.5 Flash

Output token efficiency

17% fewer output tokens

Lower cost per completed task

DeepSWE token reduction

Up to 65% fewer tokens

Much better for coding agents

Input price

$1.50 / 1M tokens

Same input price as 3.5 Flash

Output price

$7.50 / 1M tokens

Cheaper than 3.5 Flash output

This is exactly why Google is pushing it as an agentic model.

If your AI agent needs 10 rounds to finish a task, Gemini 3.6 Flash should spend fewer tokens while giving stronger answers than Gemini 3.5 Flash.

3. Benchmark Results: Where 3.6 Flash Improves

Here's the clean comparison against Gemini 3.5 Flash:

Benchmark

Gemini 3.6 Flash

Gemini 3.5 Flash

What it means

DeepSWE

49%

37%

Better at software engineering tasks

MLE Bench

63.9%

49.7%

Better at machine learning engineering

OSWorld-Verified

83.0%

78.4%

Better at computer-use tasks

GDPval-AA v2

1421

1349

Better at knowledge work

For normal users, Gemini 3.6 Flash should be better when the task needs reasoning across messy information.

That could mean reading a long financial report, moving code from one framework to another, finding problems in a repo, or extracting useful details from charts, PDFs, screenshots, and transcripts.

4. Where I Would Actually Use Gemini 3.6 Flash

I'd use Gemini 3.6 Flash as the main "do the hard work" model in a Gemini workflow. Some practical examples:

Use case

Why 3.6 Flash fits

Financial data analysis

It can read tables, reports, transcripts, and summarize key changes

Code migration

Better DeepSWE and MLE Bench scores make it stronger for engineering work

Multimodal document parsing

Useful when your input has text, charts, screenshots, or mixed file types

Internal research agents

Stronger knowledge-work score helps with multi-step analysis

Product and design workflows

Customers like Figma, Harvey, Hebbia, and JetBrains are already cited around this release

Here's a prompt I'd actually use with it:

I’m going to give you a long business document with tables, charts, and notes.

Your job is to analyze it like a senior analyst. First, summarize the 5 most important findings. Then explain what changed, what risks I should pay attention to, and what follow-up questions I should ask.

Keep the answer practical. If a number is important, mention where it appears and why it matters.
the-workhorse-of-the-new-gemini-models1

This document is 285 pages long, and it takes about 30 seconds to generate a result.

For reference: Gemini 3.6 Flash is the safer pick when the task has enough complexity that a weaker model might need extra retries to get it right.

5. Safety and Availability

Google also added stronger Frontier Safety safeguards to Gemini 3.6 Flash.

That means the model is designed to reduce risky outputs in areas such as CBRN and cyber-offense misuse (CBRN means chemical, biological, radiological, and nuclear risk areas.)

Cyber-offense misuse means prompts that try to push the model toward harmful hacking behavior.

It's live now through: Gemini API via AI Studio, Android Studio, Google Antigravity, Gemini Enterprise, and the Gemini app.

the-workhorse-of-the-new-gemini-model

6. My Take on Gemini 3.6 Flash

Gemini 3.6 Flash is the safest default choice from this release.

It gives you stronger coding results, better knowledge work, better computer-use performance, and lower output cost than Gemini 3.5 Flash.

So use Gemini 3.6 Flash when quality still matters, but you also care about speed and agent cost.

Learn How to Make AI Work For You!

Transform your AI skills with the AI Fire Academy Premium Plan - FREE for 14 days! Gain instant access to 700+ AI workflows, advanced tutorials, exclusive case studies and unbeatable discounts. No risks, cancel anytime.

Start Your Free Trial Today >>

II. Gemini 3.5 Flash-Lite: The Speed & Cost Option

Gemini 3.5 Flash-Lite is the budget model in this batch of new Gemini models.

This is the one you choose when you care more about throughput, cost-per-token, and volume than getting the strongest answer every time.

Gemini 3.5 Flash-Lite is built for fast, repeated work at scale.

1. What Gemini 3.5 Flash-Lite Is Built For

Gemini 3.5 Flash-Lite is Google’s fastest and most cost-effective 3.5-class model. It is made for workflows where the model needs to run again and again:

  • agentic search, document processing, extraction, classification, and quick computer-use steps

  • smaller background tasks inside larger agent workflows

That makes it useful for teams building AI agents, internal tools, research systems, customer support workflows, and document pipelines.

The key number here is speed. According to Artificial Analysis, Gemini 3.5 Flash-Lite can reach 463.7 output tokens per second.

That speed matters when you run hundreds or thousands of small tasks. A slower model may feel fine in one chat, but in a production workflow, latency stacks up fast.

2. Pricing: Why Flash-Lite Is So Cheap to Run

Gemini 3.5 Flash-Lite is much cheaper than Gemini 3.6 Flash.

Model

Input price

Output price

Best use

Gemini 3.5 Flash-Lite

$0.30 / 1M tokens

$2.50 / 1M tokens

High-volume tasks

Gemini 3.6 Flash

$1.50 / 1M tokens

$7.50 / 1M tokens

Better reasoning and harder tasks

This is the main reason Flash-Lite exists.

If your workflow sends a lot of small prompts, checks many documents, or runs many agents at the same time, the price gap becomes very real.

→ Use Flash-Lite when the task is simple enough that paying for a stronger model would waste money.

3. Benchmark Results: Better than the Old Lite Model

Flash-Lite is cheaper, but it still improves a lot over Gemini 3.1 Flash-Lite.

Benchmark

Gemini 3.5 Flash-Lite

Gemini 3.1 Flash-Lite

What it means

Terminal-Bench 2.1

54%

31%

Much better at terminal and coding-style tasks

GDM-MRCR v2

72.2%

60.1%

Better long-context handling

GDPval-AA v2

1140

642

Stronger at knowledge work

The jump on GDPval-AA v2 is the one I care about most for business users.

It means Flash-Lite should be much better at practical knowledge work than the older Lite model. That includes reading internal docs, checking reports, pulling useful facts from long files, and helping with simple research flows.

4. Surprising Part: It Beats Gemini 3 Flash in Some Areas

Even though it is the cheaper model, Gemini 3.5 Flash-Lite beats Gemini 3 Flash on two useful benchmarks:

Benchmark

Gemini 3.5 Flash-Lite

Gemini 3 Flash

SWE-Bench Pro

54.2%

49.6%

OSWorld-Verified

74.0%

65.1%

That does not mean Flash-Lite is the best coding model here. Gemini 3.6 Flash is still the better pick for more serious engineering work.

But it does mean Flash-Lite is strong enough for a lot of smaller coding, terminal, and computer-use tasks where speed matters.

5. Adjustable Thinking Levels

One useful part of Gemini 3.5 Flash-Lite is that you can adjust its thinking level.

  • For simple, high-volume work, you can keep thinking at minimal or low. That helps keep the model fast and cheap.

  • For harder subagent tasks, you can raise the thinking level so Flash-Lite spends more effort on multi-step work.

gemini-3-5-flash-lite-the-speed-and-cost-option-3

This gives you more control over cost and quality.

A simple setup would look like this:

Task type

Thinking level

Classification, extraction, quick search

Minimal or low

Multi-step subagent work

Medium or higher

III. Gemini 3.6 Flash vs. Gemini 3.5 Flash-Lite

Among the new Gemini models, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are the two most people can actually use. Flash Cyber is restricted, so for developers, teams, and normal builders, this is the comparison that matters.

-> The simple answer: use Gemini 3.6 Flash when the task needs judgment. Use Gemini 3.5 Flash-Lite when the task needs scale.

Choose this model

When your task looks like this

Gemini 3.6 Flash

The task is complex, messy, expensive to get wrong, or needs stronger reasoning

Gemini 3.5 Flash-Lite

The task is repetitive, high-volume, speed-sensitive, or cheap enough to run many times

That is the easiest way to avoid overpaying.

The strongest setup is using both together. Here is how that can look in a real workflow:

Workflow step

Best model

Understand the goal

Gemini 3.6 Flash

Break the task into smaller jobs

Gemini 3.6 Flash

Generate many drafts, searches, or extractions

Gemini 3.5 Flash-Lite

Filter weak outputs

Gemini 3.5 Flash-Lite

Review, refine, and make the final decision

Gemini 3.6 Flash

This pairing makes sense because the expensive part should be the thinking.

IV. Gemini 3.5 Flash Cyber: The Restricted One

Gemini 3.5 Flash Cyber is the hardest model to judge in this release because most people cannot use it.

This is one of the new Gemini models, but it is not a normal public chat model. It is a specialized version of Gemini 3.5 Flash, fine-tuned for cybersecurity work.

-> Its job is simple: find and fix security vulnerabilities in code.

Area

What it means

Base model

Fine-tuned from Gemini 3.5 Flash

Main use

Vulnerability discovery and repair

Deployment

Runs inside CodeMender

Access

Governments and trusted partners only

Public testing

Not available right now

The important part is how Google deploys it.

Gemini 3.5 Flash Cyber runs inside CodeMender, Google’s AI code-security agent. Multiple Flash Cyber agents can work together, inspect code, search for weak points, and produce one vulnerability report.

That makes it very different from Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.

Those two models are built for broad usage. Flash Cyber is built for a narrow, high-risk job.

From a security point of view, that makes sense. A model that can find and fix vulnerabilities at a high level is useful, but it is also dual-use. In the wrong hands, the same skill can create real damage.

→ My take: Gemini 3.5 Flash Cyber is probably the most important model in this launch for security teams, but it is also the least accessible for normal developers today.

Conclusion

The new Gemini models are useful, but this still feels like a Flash-series refresh before Google’s bigger move.

Gemini 3.5 Pro is still not public, and Gemini 4 is already further ahead in Google’s pipeline.

So my final take is simple:

→ Use Gemini 3.6 Flash as your main model. Use Gemini 3.5 Flash-Lite when you need speed and scale. Keep Gemini 3.5 Flash Cyber on your watchlist, because it may matter a lot for security teams if Google expands access later.

If you are interested in other topics and how AI is transforming different aspects of our lives or even in making money using AI with more detailed, step-by-step guidance, you can find our other articles here:

Reply

or to participate.