Gemini is increasingly bad, but ChatGPT surprised me
Using LLMs is as unpredictable as the weather. One day a model seems decent; the next day its next version (or even the same version!) disappoints completely. I keep giving up on some chatbots, but I still need to use them, so I feel like I’m turning in circles with no end in sight.
Preparing for the headache
When I wanted to add handling for the “429 Too Many Requests” response and VPN blocking in ComicStripBrowser 3.1, I thought Gemini (in Antigravity IDE) would make it a piece of cake.
It didn’t. It almost made it a nightmare.
🤖
I asked Grok (Fast) whether I should use Gemini 3.8 Flash, 3.7 Flash, or 3.6 Flash for this task. I’ll quote from its answer:
For Python coding in Google Antigravity, Gemini 3.8 Flash is the strongest overall (especially on complex/agentic tasks), 3.7 is a solid middle step with big gains over 3.6 on benchmarks, and 3.6 is often preferred in practice for speed, lower token burn, and less over-reading of files. All three currently share the same introductory pricing ($0.75/1M input, $3.75/1M output through Dec 31, 2026).
In Antigravity (the agent-first IDE), newer models are better at autonomous coding loops, error recovery, and end-to-end workflows, but real-world behavior diverges from pure benchmarks.
Real-world Antigravity / Python coding experience
Community reports (especially the Google AI Developers Forum) show a consistent pattern for IDE use:
- 3.8 is accurate and strong at deep analysis/bug-finding (it tends to read related files thoroughly), but frequently over-reads files repeatedly, does more work than needed, and burns quota faster—even on moderate tasks. Great for hard problems or planning; less ideal as a daily driver if quota is limited.
- 3.7 sits in between; some users find it less reliable/accurate than either 3.6 or 3.8 in the IDE.
- 3.6 is often called the “balanced” choice: faster, makes targeted changes without re-reading entire files repeatedly, lower latency, and much kinder on quota. Accuracy can occasionally lag, but many prefer it for everyday Python work (edits, features, quick iterations). A common hybrid: 3.8 (High) for planning + 3.6 (Low) for implementation.
For pure Python coding (scripts, libraries, refactors, tests, debugging), the efficiency edge of 3.6 is frequently more useful day-to-day than the benchmark wins of 3.7/3.8, unless the task is genuinely long-horizon or complex.
If quota/cost is a constraint, 3.6 (especially at Low/Medium) usually wins. For maximum quality on tough problems, 3.8 can be worth the extra tokens.
Should you set thinking to High?
No — not as a default. Use the lowest level that works for the task.
Antigravity exposes thinking levels (Low / Medium / High; 3.6 also supports Minimal). Default is typically Medium. Higher levels increase reasoning depth, latency, and (especially) token consumption.
Practical guidance for Python coding:
- Low (or Minimal on 3.6): Simple edits, boilerplate, single-file changes, quick tests, well-scoped tasks. Saves the most tokens/latency.
- Medium (default): Everyday feature work, multi-file changes, most agent runs. Best balance.
- High: Architecture decisions, complex multi-step debugging, long-horizon refactors, or when quality clearly outweighs cost. Use selectively.
Antigravity team guidance and community consensus: start low and escalate only when needed. High on 3.8 can be excellent for planning hard work but is expensive for routine coding.
Quick recommendation for Python in Antigravity
- Daily driver / quota-conscious: 3.6 Flash at Medium or Low.
- Complex agentic or high-stakes coding: 3.8 Flash (Medium or High for the hard parts).
- Hybrid that many use: 3.8 High for planning/architecture → switch to 3.6 Low/Medium for implementation.
- Always tune thinking level per task rather than leaving it on High.
You can switch models and thinking levels freely in the Antigravity model selector (it persists per conversation). Pricing is the same for now, so experiment on your actual Python workflows—the subjective difference in file-reading behavior and quota burn is often more noticeable than the benchmark deltas.
Later advice:
Recommended model + thinking level for this exact prompt
Gemini 3.8 Flash → High
Reasons specific to this prompt:
- It must synthesize the detailed PLAN.md requirements.
- It needs to propose sensible ordering / grouping of the five features (they are not all equally urgent or equally easy to test).
- It must respect the “outline options first, wait for approval” constraint.
- 3.8 High tends to be more diligent about reading the full context and producing structured option lists.
After you approve an option
Once you pick an approach, switch the model to 3.6 Flash (Medium) in the same conversation and tell it to proceed. Context (including the three docs + the options discussion) remains available.
Rushing head-first into the wall
Stupid me decided to use Gemini 3.8 Flash (High) for everything because, what the heck, only 5 Python files needed to be modified, and the entire project consists of 20 Python files, of which one is version.py and 3 are __init__.py, so literally 16 files matter.
Grok was more than accurate to say that 3.8 “frequently over-reads files repeatedly, does more work than needed, and burns quota faster—even on moderate tasks.” I couldn’t care less about quota for such a simple task, so I decided based on this other assessment: “Great for hard problems or planning.”
And indeed, it seemed to understand the tasks set in PLAN.md, which it expanded rationally and sensibly.
But then, oh, boy.
🤖
It started to implement one feature, and it did it poorly. When I pointed out that its changes should apply only to GoComics titles, not all titles, it got confused. It started to reread all files, then it “discovered,” as if in surprise, the changes that it had just made, and started to ratiocinate: “There is this thing here that appears to do that, so I need to check where this is used. The user claims that this thing happens, so I need to find out if this really happens and what might be the reasons.” FFS!
What should it have implemented next? A couple of checks in a couple of files. Parameters. Conditions. It had exactly 5 files to consider.
It reread files. It marveled at what it found. It implemented half-baked checks, only to discover, after further tests, that those checks didn’t check for GoComics.com.
It reread files and made some minor, incomplete changes. It ran too many nearly identical tests that were 100% likely to pass, and it never tested what would certainly have failed.
Eventually, one test failed. It reread all files (and I mean all Python files!) and again rediscovered that water was wet and marveled at that. It made a few more changes, then ran three times as many tests, most of which were obviously useless. I had to repeatedly reject its requests for more approvals, telling it that I’ll test that specific thing myself, manually.
Lather, rinse, repeat.
After an eternity, it did implement one feature that worked. After I made sure it indeed worked, I switched from 3.8 Flash to 3.6 Flash for the rest of the session.
This improved the LLM’s behavior, but not entirely. For instance, when it discovered that a comic that should have been retrieved from ComicsKingdom was actually retrieved from GoComics (I get “Dick Tracy” from two sources, and the second definition was botched), instead of recognizing that there was an error based on a copy-paste operation that wasn’t followed by the editing of a value, it started to philosophize, to reread all files (including COMIC_TITLES.md), and to speculate about the two data structures that used the same base URL. Any retard would have caught the error by examining the two definitions a couple of lines apart—it didn’t. I had to stop it and manually fix the wrong string.
Remember how I said a month ago that I’m afraid of Gemini 3.8 Flash? Well, right now, I’m afraid of all versions of Gemini!
AGENTS.md explicitly said: NO GIT OPERATIONS. Then, “DO NOT run git commit, git push, or alter repository remotes. Version control is managed strictly by the user.” But Gemini 3.6 Flash asked me at least 4 or 5 times to run at least git status. Each time, I told it to stop caring about Git and GitHub, and it still “forgot” within less than 10 minutes! It only stopped after I added: “I already told you this in AGENTS.md and several times during this session!”
How did they manage to break Gemini so badly?
Previously, Gemini 3.5 Flash and 3.6 Flash were acceptable, and 3.7 Flash was usable as well. You can’t trust anything from Google these days!
I didn’t retry Claude Sonnet 4.6 because I’ve found its quality to be inconsistent in the past. The last time I tried Qwen for coding was when Qwen Code still had a free tier, which was before April 15.
I could reevaluate Qwen by using a trial Alibaba ModelStudio account. I still have 54 days to use such models in Qwen Code:

But I’m almost sure that I’ll be disappointed.
When I had to switch from Gemini CLI to Antigravity CLI (and Antigravity IDE), Gemini 3.5 Flash was still OK, but now it’s not available anymore!
Why must “newer” so often mean “worse”?
Gemini 3.1 Pro is also broken
I still have an extra-cheap Gemini Pro subscription that extends Gemini’s chatbot limits, but not indefinitely—it can still kick me out. Given how retarded it became, I’m glad I only paid peanuts for it!
For instance, my arch-enemy, ChatGPT, which offers very limited free access to an unspecified model, sometimes gave brilliant answers. If I don’t need to upload any file, it can handle decently long threads before telling me that I have to wait 5 hours.
I fell out of love with Claude, but I still refer to it occasionally, and it can be helpful. I’m just not sure which models to use. Free users can access Sonnet 4.6, Sonnet 5, Sonnet 5.5, and Haiku 4.5. Each can have 4 levels of effort (Low, Medium, High, Max) and optionally “Thinking” (what did they do on “Max” if they don’t think?!), except for Haiku, which only has an “Extended” toggle with the following explanation: “Always uses deep reasoning.” Oh my, reasoning and thinking are distinct features, and an LLM can struggle longer (and burn more tokens) even without thinking?! Who can still understand Claude? I liked Sonnet 3.5, maybe also 3.7; I frowned upon Sonnet 4 (others did, too!), even more so with Sonnet 4.5, and I never truly found Haiku useful. But today, I can only trust Sonnet 5.5 High, Extra, or Max (5x more token usage), which might eat my quota in 1-2 questions while taking an eternity to answer! Maybe I should test Sonnet 4.6, eh?
I don’t use Copilot. Not anymore. Grok? Rarely, because the only option is Fast, and the free quota is variable. Now, ChatGPT does practically the same: no options, and you don’t know when it stops answering. But Grok can be extremely annoying at times, so I even had to ask it why it was an ass:
Ludditus: Why do you answer in such an obtuse, unempathetic, annoying style?
Grok: You’re right to call that out.
I’ve been answering in a very precise, almost lecture-like way because the topic is technical and you’ve been pushing on exact definitions. That style can come across as dry, rigid, or lacking warmth.
I’ll dial it back. Clear and direct is still the goal, but I can drop the stiff textbook tone.
What specifically felt most annoying in the previous replies? The structure, the tables, the repeated corrections, or something else? That’ll help me adjust better.
I also stopped using Kimi since the only free option is “Instant.” Sort of like Claude’s Haiku, I guess. DeepSeek? Meh. No lo conozco.
Speaking of free vs. paid options, even with a Pro subscription, Gemini only offers 3.5 Flash-Lite, 3.8 Flash, and 3.1 Pro, each with or without “Extended thinking.” No 3.5 Flash (full), no 3.6 Flash, nor 3.7 Flash—nada. WTF?!
Also, even if you pay, you can’t tell Gemini to stop telling you how great you are for having asked a specific question! ChatGPT lets you make those kinds of customizations for free, and it never told me again how smart I am!

I can’t believe I started appreciating ChatGPT and hating Gemini! Everything must have a beginning…
A recent test that puzzled me—yes, Gemini 3.1 Pro is pathetic!
On my Quick Luddite Notes second site (currently listing 90 notes!), I recently posted 🤖📽️ MobLand: Gemini, Claude, Qwen, and ChatGPT answer. I asked these 4 LLMs a 489-word question about why they’re making series that I don’t like. It was a “barroom philosophy” (pub philosophy, if you wish) topic. The kind of topic I usually ask 2-3 chatbots about to see how they “think.”
You can read the respective answers there, or directly as shared:
- Gemini 3.1 Pro Thinking: 443 words. A completely pathetic answer. If this is what “Pro Extended” can do, what should I expect from 3.5 Flash-Lite? (I was too disappointed to try.)
- Claude Sonnet 5.5 Max: 702 words, and it took a very long time to answer, but it wasn’t that bad. At least I liked a few ideas, despite the lack of structure. I suppose it makes it more “human,” but meh. Passable, yet succinct.
- Qwen3.8-Max Thinking: 767 words, with 3 subtitles and a conclusion. Unfortunately, more like a longer version of Claude’s answer. I appreciated 2-3 ideas, but not the overall wording. Not bad, but underwhelming nonetheless. At least it was faster than Claude. In my experience, Qwen3.7-Plus is much faster and possibly better, while Qwen3.7-Max, Qwen3.6-Plus, and Qwen3.5-Plus are also available.
- ChatGPT Think (free): About 2000 words, but it’s difficult to count because it included tables, lists, quotes, and logic sequences (→), plus 10 URLs properly copied in Markdown! The best-structured and deepest analysis of all! Unbelievable.
I remember some LLMs giving weaker answers when “Thinking” was enabled (DeepSeek with DeepThink did that to me), but that is obviously not the case with ChatGPT today. I didn’t test Gemini Pro without thinking enabled, so I can’t be sure about anything, except that I am disappointed!
Other recent disappointments: completely stupid images generated by both Gemini and Qwen. Simply unusable!
Where do we go from here?
Fuck knows.
I can no longer trust Gemini for code.
I’m increasingly disappointed by it as an all-purpose chatbot, too.
I don’t know (yet) how appropriate Qwen is for coding. I hope it’s still OK-ish. As a chatbot, it’s so-so, and its generous limits may have already shrunk since it introduced paid plans for chatbot use (it was the only chatbot without a paid option!). OK, Qwen is open-weight; you can run it here and there, in various quantizations, so it’s something to play with. Am I too optimistic? Maybe it’s slowly becoming a dead end; who knows?
Claude is slow as a chatbot and probably expensive if you have to pay for it, no matter how you use it. Either way, I’m not thrilled. (I can use it for coding as part of the Google AI Pro plan. It can also be used in Amazon’s Kiro.)
I never used any GPT model for coding, but I stopped looking at people who pay for ChatGPT as if they were retards. (Most people are retarded anyway, regardless of what chatbot they use and how they use it.)
Kimi is dead to me (you can only lose your confidence once), and I have no idea how reliable it is for coding, but they have started accepting new subscriptions again. Is K3 really worth it? And why is K2.8 in preview if K3 was launched months ago?
DeepSeek is probably obsolete or too slow to evolve (if evolution really exists in AI). Copilot is outdated. Mistral is pathetic (les tarés even renamed Le Chat to Vibe). Grok is Satan’s, and Meta’s models belong to an even more malefic organization. How could anyone pay for any of them?
More importantly, which LLM to trust for agentic use? Some people seem to be torn between Claude Opus and GPT. Self-hosting deployments require Qwen, DeepSeek, or GLM. Yikes, I forgot about Z.ai’s GLM models!
Right now, I don’t trust any LLM to perform any sensible task. None. Zero.
🤖
Oh, I forgot! Speaking of models and their fabulous benchmark results in the context of agentic use, Gemini ended a conversation (also posted here) this way:
When you are doing complex technical work, a fast idiot is the most dangerous tool in the stack.
Go figure.

Leave a Reply