Wait, did you let the same agent write the tests that wrote the implementation? Or was it that the plan was the wrong way round, and both the generator of code and the generator of the tests followed that same incorrect plan?
Also, what were your mechanisms for reviewing the plan against the described business outcome? Anything else you could have done there to catch the backwardness?
And finally, did you run any of it across models from different foundation labs? I'll often run important stuff generated by codex across Anthropic, grok, and Gemini with an opus or fable judge...
First of all the necessary context: this was a hobby project, and I am not a professional software engineer.
In this particular instance I'm not sure what happened. I don't know where it got the idea from to do it backwards.
My best guess is that the to-do items were too vague. I built up a lot of context in my head from back and forth with the agents, and so my mental model was pretty solid, but probably an insufficient amount of that was encoded in the repo itself.
I'm not sure what you mean by reviewing the plan — do you mean making a detailed implementation plan before beginning the work?
Most of the changes were pretty straightforward, or at least we'd already worked out most of the details and put them into to-dos. So for the most part my prompt was just "alright go ahead and implement the next thing on the list."
If I had to guess I'd say the main reason it went wrong was because the why was missing. The to-do item specified what work remained to be done, but it did not explain the reason for each item. I think that was the core of the issue.
So it probably ended up seeing a bunch of individual changes out of context and not understanding what the point was supposed to be.
---
>did you run any of it across models from different foundation labs?
Working on a different part of the same repo a week later, I used Fable and Sol to find issues. Then I let them run cross critique on each other's reports. Then I had each of them generate a plan, and then I had them do cross critique on the plans. And then I repeated that until they were both satisfied with the results, i.e. until the two plans merged into one coherent plan.
It was interesting because they're both had different strengths and weaknesses in different parts of the process. (One of them sound more issues in the initial phase, but the other came up with a more comprehensive fix for each one.)
I've seen very promising results for model alloys with security research (it was posted here about a year ago[0]), so I wanted to give it a try myself.
There doesn't seem to be a very convenient way to do it (maybe one of the new harnesses which runs the proprietary harnesses as subprocesses?), so I just passed markdown files between the two models manually.
That's what my microeconomics professor said shortly after Citizens United. "If political campaigns were that good of an investment, they'd be investing more!" ...and then they did. Outside spending absolutely took off after 2010 and shows no signs of slowing down. The conclusion he jumped to was false and it was false because he ignored the possibility that price discovery might take a while.
Yeah, this has been my biggest contention with the models everywhere paradigm. I will definitely use a supervisor pattern and advisor pattern for simpler models when I just want to throw something together interactively with Fable and a few sub-agents. But for anything I am doing repeatedly, I built a deterministic orchestrator to run the steps and I have a hard rule that the skills that run from my repo can only ever be a thin shims calling to my central system so that rather than having agents enrich a person or call an API or do research they fire off predefined multi-step deterministic plays that just might have models for classification generation summarization and/or review.
In addition to the reliability and cost benefits you also then get shared capacity broker capabilities so if you're only able to enrich so many people or only capable of doing so many CI runs if that's all done deterministically you can have some intelligent orchestration in your main server to manage that limited capacity across all your agents that are trying to use it at the same time.
Also have a workspace as a personal email and ended up getting a personal gmail just to try out the subscriptions before I gave up.
I have multiple anthropic and OpenAI max plans. For Gemini I just use my Cursor $200 a month plan (which also gives me the ability to try grok, conductor, etc)
I would take the other side of this bet. While I agree that the impact of any given advance is likely to resemble a sigmoid curve, I think there is a material chance of "stacking sigmoids" creating something that looks exponential.
To take a simple example, look at the progress of technology over the last ~500 years - it seems to me that the rate of change continues to accelerate despite many of the logistic curves flattening along the way.
There are still huge unanswered questions about whether or not the stacking sigmoids will favor the incumbents. But I would not definitively bet against the people with the most compute data, talent and money.
I'm unsure what exactly you mean by sigmoid stacking.
I'd argue that a lot of important technologies (like circuit design!) started at basically zero in the last century, and the progress was actually exponential for a significant time (=> because we started so low on the curve).
But if you reduce things to a single metric (wealth per person? total energy available to humanity? global industrial/construction output in tons?) I can't think of anything where such "stacking" successful subverted the sigmoid trend (or looks like it will, long-term).
AI can materially improve PRs - you just have to do it right.
Firstly there are not going to be a material number of non-AI patches written in the future. I know some people still ride horses, but when compared to the 17th century the percentage of travel primarily enabled by literal horse power is way down. Same way goes "artisinal, hand crafted code"
Secondly the problem isn't the AI - it's the crappy prompts and lack of intermediate artifacts and quality gates (deterministic and adversarial) in the generation pipelines.
Thirdly, don't disallow AI (I mean, feel free, you're doing something without charging for it - do you) - disallow crappy PRs, verbose descriptions. and bloated code that doesn't fit your coding standards.
Fourthly, ship your standards. Document what a good PR looks like, the comments and examples, the search for open PRs to ensure it isn't a dupe or a won't fix. Then ship your coding standards - have some LLMs infer them from your current code base and review (bonus - anything it infers correctly that you don't agree with, get it to re-code to your new standards).
Then set up a PR reviewer - start with deternministic gates for the format of the PR and some broad heuristics, then run it through adversarial reviews. Anything that looks great, feel free to run by a human either before the merge or after just to keep an eye on things.
The deterministic gates will cheaply get rid of 95% of the slop (AI and human - I've seem plenty of human slop over the years) so you can focus tokens on the good stuff.
I think "no-AI" policies are really just ways of gatekeeping low effort contributions from people who have no idea what they're doing. This is a good thing!
To use the current example of Godot--game engine development is really, really, __really__ hard. Building on your engine in a way that doesn't lead to performance regressions, introduce subtle bugs, etc. takes a lot of know how.
If you care about quality and performance, then rejecting work from people who aren't capable of understanding how their contributions impact the greater engine as a whole is a logical choice.
The problem isn't the AI - it's the assessments. It used to be that shipping a large, well-researched essay with multiple citations was proof of work. Now it's proof of prompting, not learning or effort.
If you want to test for memorization, you need a live test to do so.
If you want to test for comprehension and understanding, require sessions with a tutoring AI and grade the level of comprehension exhibited.
And then please, please test for prompting with challenging assignments that would usually be beyond a student's skills and make sure they know how to drive the machines to do valued work.
I totally agree that there is utility in memorization and comprehension. I also know that by the time students graduate, AI will be the job and if they can't pair successfully with agentic workflows they will not be much use in the work force for many roles.
I auto tune my prompts to a locked model version based on production data used as evals with holdback data. I think the use case for this would be one off interactive prompts? For now I just run those all against an Opus 4.8 MAX and I'm sure I could downtune, although for interactive my opening prompt isn't always reflective of my overall goals for the multi turn session.
I'm just trying to figure out why on the fly routing would beat testing and tuning and locking models and versions for each class of call, with evals and auto tunes running to explore more possible models for commonly run classes of prompt over time . . .
"Based on your subscription tier and local hardware here's a list of models that fit and process definitions your biggest brain will comfortably handle."
I guess that sounds a lot like moving your evals and auto tunes to a third-party, but I don't have the time, budget, or inclination to create a system like this out of whole cloth and then keep it relevant.
I could see something that provides on-the-fly routing information being useful, but actual decision-making is too dependent on context.
Honestly my goal is to learn how to teach an agent to build a maintainable product, so I'm way more interested in the learnings at the agentic level (how to prompt/direct/manage context/restrict tool use, provide reusable shims, etc) than getting into the details of a css bug. That's just not a level of abstraction with sufficient leverage for what I'm trying to do.
I stopped coding a while back because I could have more impact directing a team of developers than writing code personally.
For my use case, the agents are now how I can have that scaled impact.
Absolutely. All of these "but you could have done that easily" from frontend developers or backend developers or systems engineers -- like yea, if I have the time or interest in those things, sure. But I don't. I care about an end product way way more. Blows my mind that there are legions of people building things that they don't think are important enough to get to the finish line quickly and efficiently.
Also, what were your mechanisms for reviewing the plan against the described business outcome? Anything else you could have done there to catch the backwardness?
And finally, did you run any of it across models from different foundation labs? I'll often run important stuff generated by codex across Anthropic, grok, and Gemini with an opus or fable judge...