Hacker Newsnew | past | comments | ask | show | jobs | submit | galkk's commentslogin

Opened this link to quote that. The level of reasoning is "trust me bro". Even if some things must be simplified for uneducated reader, the quote from above should be embarassing for author to write

Our database 12.34 was working great, but 12.35 deployment had some optimizer changes that had regression on exact scenario that you have in your statistics. Shit happens, sorry. Use this hint.

I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.

For information retrieval and collation related tasks (reading lists, deep dives on subjects, etc.), Gemini is way better w.r.t. other models in my experience.

This difference is probably due to Google's web knowledge and free pass to YouTube. However, I'm happy what I got from it so far.

When I ask the question once in a blue moon, it can generally one-shot the answer, even.


I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.

This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. Which, to be fair, is the case with the free models.


> I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.

I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Google/GMail) for anything that is not coding.

I find Gemini better/quicker/more polished for basically every single subject out there that is not "write me lines of code".


Same here.

I use gemini for everything not-coding, from doing research, to have custom personas for more niche topics (and feeding more detailed knowledge in these cases).

For coding and image editing, right now I find ChatGPT superior. And for software architecture designs or planning Claude is the best since a while. I still have to try Grok to be fair.


To be fair, it has been about six months since I tried a paid Google model. I will give it another go to compare. Hallucinations were the main issue back then but perhaps it has come a long way.

Much respect for being able to admit that you haven’t used a model in a while after saying folks probably haven’t used the models you use.

Okay I just tested 3.8 Flash (High) on some real world problems I have used Opus (High) and Sol (high) to solve. The outcome here is terrible.

One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc.

3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info.

I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy.

But don't take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt...)

This has been a useful exercise. It's important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.


Gemini's search harness in the Google app is (ironically) bad so it makes the model look bad.

If you really want to compare apples to apples you need to test Gemini models against other models using the same third party search harness.

Otherwise you are largely measuring how much computation the model provider is allocating to a search harness.


I used [AI Studio.](https://aistudio.google.com/) It's possible the same issues are present there, but AI Studio is intended for serious work. I don't see why they would intentionally hobble the model's capabilities in AI Studio.

I've got the 100 paid to all 3, gemini and chatgpt are a level above claude for a lot of my work now. claude i actually fight with if its not just write code.

> you're just not using the latest model, bro

Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world.

P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.


> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.

The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.


The tasks you actually need to do trump benchmarks, yes. I haven't tried out the ridiculously expensive models besides the latest Gemini, and it gave from equal to slightly worse results than latest DeepSeek, at a far higher price.

It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.


I think 3.8 Flash is on par with DeepSeek on some benchmarks and tasks (not coding or design), but it's not close to Sol/Astra or Opus/Fable. I would not consider a $20 subscription "ridiculously expensive," but I suppose that is a relative term.

I would run into the use limits very quickly, and (for Anthropic) have to switch frameworks.

By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.


No argument that DeepSeek is much cheaper, but you certainly get what you pay for.

artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.

Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.


It's very impressive that DeepSeek 4.1 beats Astra in one benchmark, but I presume you are aware that Astra wins in almost all other benchmarks? These are some of them: https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-...

Where did he claim otherwise?

Take your meds.


Impressive how feverishly you defend the steaming pile of shit that Gemini is.

I love it as a variation from the others. 3.8 flash is the best back-and-forth model for iterating imo, but would not use for long horizon

As I mentioned - I find it awful for iterating. My recent example: I was researching shoes for toddler, wide, boa or similar mechanism instead of laces/velcro. I did same prompt in ChatGPT and flash. After several back and forth responses/clarifications flash completely lost track of what I’m looking for and started suggesting nonsense.

I see! I've only really used it for coding. You should design velcrobench!

I also had that issue happen to me, surprisingly only when I used Polish, not English. But other than that I actually like Gemini, I check stuff against it all the time, especially when walking my dog. It also works in AndroidAuto for me, but I use very basic stuff like changing Spotify music. I was on iOS before and there's no comparison to old Siri that was just garbage. I use Gemini practically every day, it's fine for the most part IMO. They really do need to polish integrations though, app connectors barely every work outside of Google's own apps. Using Oppo's app Mind Place through Gemini f.e. is just bad experience and almost never works. For example, I had pleasant experience of Gemini finding me places for a walk/hike on vacation in Tirol, where I specifically required not too much of an ascent and an asphalt road for a stroller.

I wonder how much life DeepMind has left in it, especially after Hassabis's departure. Google execs must be having discussions about simply throwing their weight behind Anthropic since they already own so much of the company.

They do a lot of weird things with the context in their user facing products like Gemini. Google seems to always have trouble with their harnesses

Have you tried Gemini 3.8 Flash recently ?

I was like you before, Gemini was the worst model to me.

Then 3.8 came out. At first I was sceptical, but this model *is* able to do useful things ! Complex things.

Of course, it is NOT perfect. But for things like small/medium complex tasks subagents, it's perfect.

Now, is it worth the money vs Astra ? I don't think so. But still my point remain relevant.


I agree. Lot's of people really like it. For me it often just forgets all context and starts showing random slop. It's super clear as, when I ask it what happened to some element earlier in the conversation it tells me it does not have that. It might be good if it told me, but randomly lose the plot is quite frustrating.

Claude does it occasionally but it's a more a soft landing earlier context seems to be compacted, not completely lose the plot.

I just cancelled my pro subscription. I really wanted it to be good but not yet.


Definitely true in some cases, but I find Gemini to be less biased about certain topics - which is quite nice, especially when I'm just trying to hear the facts.

I’ve found this to be true as well. It’s got really poor attention and will derail into world-building fast. I’m convinced Google is just shipping it to capture market share but they know Gemini isn’t ready for serious use.

i have programmed a dozen scripts with google search AI lol

they'll live in my rc for decades


Ah yup for quick one-off one-liners or small Bash script, Gemini works perfectly fine too. For longer scripts I use another model.

It is by far the least accurate of any model I have used too. For a company that was started to organize the world's information, it has by far the most misinformation I've encountered.

I couldn't even get it to tell me how to pay for antigravity, it sent me on some fruitless paths and eventually said "you shouldn't pay for this, it's too hard to figure it out."


I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.

I would like to see chat logs etc and understand how much of a progress was done by human.



I hate when they measure the size of source code instead of the size of a binary.

I appreciate .kkrieger much more than this monstrosity


Time for bitcoin classic++?

Are you thinking of ETH? All the bitcoin classic forks are over various aspect of network rules (eg. block size or block reward), not to roll back a transaction like ETH classic.

You’re right, I misremembered. Yeah, I was speaking about ETH Classic thing. Thanks for pointing that out.

Yeah… because in 15 years bitcoin has never forked for a hack. What a brain dead comment

Idk, I never cat problems with getting a random Home Depot employee to cut it, the problem is that they are usually very rough cuts, far from square and precision


To me the book that changed the way I think is “pragmatic programmer”. I haven’t re-read it in a while, but for a long time it was my first recommendation.

——-

And I was recently thinking about books that my opinion has significantly changed about the longer I’m in my career, I realized that “Peopleware” truly sucks. It is such feel good fairy tale, that actually harms the reader. The real world is much different and much more cynical.


> So what makes up the remaining 40%? Under the Bundesförderung für effiziente Gebäude, households receive 30 to 70% of eligible costs. If an installer raises a €25,000 quote to €30,000 and the household receives a 50% subsidy rate, the household’s net cost rises by €2,500 while the installer receives €5,000. This means the pressure to negotiate a better deal is halved because of the way the policy is designed.

I’m sure that well designed kickback scheme is hiding somewhere here.


I'm not sure about "well designed".

Rather, my thought is that such a subsidy encourages sellers to raise the price of the unit while including something else along with the unit. Suppose they raise the price by that €5,000 and include somewhere €3,000 of something else in the deal.

Installer pockets €2,000, customer pockets €500, government pays.

This can be emergent behavior, not necessarily designed.

We see this with Medicaid (healthcare for poor people) here in the US. Each state runs it's own program, but the federal government pays a large *percentage* of the cost. Enter drug companies offering the state big rebates to cover their brand name drug, states reacting by putting the brand name on the approved list and excluding the generic version. The total bill is a lot higher but the state gets back enough that it saves them money. This certainly wasn't intended in the design, but there's enough money in it that it hasn't been fixed.


> not sure about "well designed"

I'm not sure cause I'm not an English native speaker and on the Aspergers spectrum, but I think this is sarcasm.


Sometimes things like that are intentionally designed to funnel payoffs to somebody.


Sure, I mean, look at Reddit...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: