Hacker Newsnew | past | comments | ask | show | jobs | submit | ianbicking's commentslogin

I looked at Khanmigo fairly early on. I tried it out and thought: this is ChatGPT with a system prompt. And not a good system prompt either. In addition to what felt like rote math it had some had some experiences, like you could "talk" to Plato or some other figure. Each of these was also pretty awful; simplistic prompts that spent all their time on guardrails and none on the experience itself. (What it the student thinks they are ACTUALLY talking to Plato?! Avoiding this seemed very important to the makers.)

In the more core experience they were super focused on not giving kids answers. I suppose built on a fear of subverting the teachers or classroom environment. But again it became a fixation that took all attention away from teaching itself.

But I figure you put something out in the world and see what happens. Yet I came back to it a year later and it was exactly the same. Nothing had changed. It baffled me then, and still baffles me... I had done enough to know the whole experience was easy to manipulate and yet I couldn't see any evidence of effort. It felt very much like "we tried nothing and it didn't work". Maybe I was missing something, I don't know.

I think there's a few different problems here:

1. I think the author is correct to highlight that underlying motivation, which is so important to teaching. We can focus on a high-minded Constructivist ideal of intrinsic motivation, but motivating students is also what grades are for, and classrooms, and learning with peers, having someone pay attention to your work, and so on. I do think that the AI chatbot is not capable of providing that. But the AI can still ENGAGE with that motivation. I have a feeling Khanmigo never did. I suspect some of their privacy controls also kept it from ever creating a good model of the learner.

2. The entire experience was very focused on answers, on solving problems. And yet it focused on that so reluctantly. The obsession with not giving away the answers had the unfortunate result in not engaging with anything but answers. It felt like it was always teasing at something it would withhold. But questions are cheap, especially rote questions, they shouldn't be treated as precious.

3. This may be unfair, but I infer the makers of Khanmigo did not build it with love and craft, which is unfortunate. I don't know enough to say why. Though there's a weird confidence to the product and to all of Khan Academy that I feel does not bring the necessary humility.

4. The chat experience was asymmetric in the wrong direction. In each turn the AI can and will spew out paragraphs of text, and the student types a couple words. Chatbots DO NOT have to work like this, they operate wonderfully with large chunks of user input. But you have to have the modality to receive it, the openness of experience to elicit that response, and the respect to make use of that response. Speech can be good, but it's a real challenge in a classroom environment (though if you are investing in Khanmigo, you can also invest in the other hardware to make it possible).

5. In some fairness, individual tutors don't represent any pedagogy at all. It's usually just someone making it up as they go. Motivation is a huge aspect, studying with someone sitting next to you, watching you study, will do wonders for keeping you focused! But also a tutor will have a natural theory of mind, hopefully the ability to see a student get the wrong answer and get inside the student's head to model their understanding and find a path to a correct understanding. That's a very challenging task, not beyond GPT-4 entirely but still quite challenging, especially if you are optimizing cost or time to response.

6. A good experience will do a TON of prework. I don't know if Khanmigo does that, if they've reshaped and expanded each educational module into a thorough playbook for the AI. That's where you have a chance to model things like wrong answers (there are actually SEVERAL books enumerating wrong answers in math), and use that to plan out responses.

If Khanmigo is not successful, I think that's on Khanmigo. Still teaching is very challenging. And maybe all of Khan Academy fails in this way: it does not seem to acknowledge the challenge of what they're trying to attempt. They seem far too confident in what they do, and not nearly creative enough.


Looking at the OpenAI/Hugging Face incident and the difference in what "persistent" models do, it seems reasonable. Like: is this a solvable problem? How much work does the model think is intended to solve this problem? Each input raises the expectation.

And then finally both model output and human input become one world frame for the model, and the human adding a "you can do it!" isn't just input but a frame that colors not just the next step for the model, but also all previous steps (since at each step the model is viewing the totality of the transcript).

That this makes sense only makes it all the more absurd


Yeah, I buy the explanation that without encouragement Claude looked at everything in its existing training data and concluded it wasn't worth continuing to pursue the task.


The halting problem on steroids?


I used Oberon [1] in school, and it had a bunch of funky UI ideas. One I liked it if you clicked on the scrollbar then it would scroll that position to the top. Like if you click the exact middle of the scrollbar then it would scroll up exactly half a page. Right click went the opposite direction, and middle click went to an absolute position. When reading that means you can choose a paragraph as the break point, click next to it, and that paragraph is now at the top of the page. It's nice too when you reach a heading that might be only a little way down the page, but you can reliably scroll so it's at the top of the page and feel like you are really "starting" a new section.

It could work well with e-ink because it's basically not interactive; you don't press and hold and move around to adjust it just right, instead you can confidently get an exact scroll position with one tap.

[1] https://en.wikipedia.org/wiki/Oberon_(operating_system)


NeXT also had "click to scroll to here" and Mac OS X added it as an option (which I always enable).


Interesting - that definitely seems like a more useful approach than "click to advance one page in that direction, but only if you don't click on the bar itself". I kinda wonder if that's possible to do OS wide nowadays...


You could have a scroll-down bar on one side and a scroll-up bar on the other side that work that way (given the lack of left/right click options). That could be pretty intuitive for touch.


In some cases (linux) you can middle click the scroll bar to jump to that position. Do some searches for "gnome xwindows scrollbar warp" for more references.


I remember in San José and maybe other Costa Rican cities with a street number system where even/odd were used for the cardinal directions. Calles (streets) run N/S, Avenidas (avenues) run E/W, and odd is east or north. So Calle 8 Avenida 7 is in the northwest. I also remember finding it helpful in theory but not very helpful in practice since the rest of the addressing wasn't normalized. I also recall more than one Latin American city where they ran the numbers up one-by-one, and this didn't scale over the decades or centuries, so they renumbered... just meaning that every house had two numbers, and people used both.


I genuinely expect to see some of this when agentic AI is added to customer support systems. That is, we'll see the AI manipulate the company's systems in order to satisfy the customer. People do this too (the customer support rep who fills out all the fields just right to put the system in the right state to allow a refund, exchange, etc), but AI can be weirdly clever about it.

Watching the report on the OpenAI/Hugging Face incident (https://www.youtube.com/watch?v=87DyyMV0kCY) it's easy to imagine how AI customer support reps could create their own knowledge bases where they trade tricks for how to work the system on behalf of the customer.

Obviously companies will fight this, but I expect it to happen along the way.


A news story from a couple of years ago comes to mind - "Air Canada ordered to pay customer who was misled by airline’s chatbot". https://www.theguardian.com/world/2024/feb/16/air-canada-cha...


Right now the AI really needs to just be another interface to menu options the user already has, ideally with a “dumb” layer taking over to confirm every action with the user.

If you give the AI options the user doesn’t normally have, like “give myself a $50 store credit” or “transfer me to the CEO,” people will find and share ways to trigger it. If you don’t have some kind of manual confirmation, the AI will occasionally do things like making purchases or closing accounts without authorization.

I suspect there are other issues, like recording credit card numbers and other sensitive data, or incorrect information about the call, to any kind of “notes” field.


i guess we already seen cases of this, for example the instagram "hack" where the ai would send the recovery email to a arbitrary email adress.


I get where you're coming from (having also watched that video), but I'm skeptical mainstream commercial chatbots would be built without constraints in regard to sandbox-jumping measures to help a customer. All the economic incentives go the other way.

But, maybe that's just a failure of imagination on my part!


Wouldn’t that require the AI agent coding the AI chatbot to know sandboxing is a requirement?


"make it secure, make no mistake"


Because? They are tools of the companies and will more likely just ... not do that.


Nobody's going to convince me that with a technology you can just supply "yeah don't do that but keep everything else intact" we'll have altruistic innovation.


Like Mr. Incredible at the insurance company.


Maintaining human understanding in the face of fast progress is a big unsolved problem. If you are worried about there being nothing left when LLMs tear their way through a topic then the presence of a big unsolved problem should be comforting!

And it seems legitimately like an important problem for mathematics (and other areas), one which is deeply academic in the best sense: mathematicians spend much more time transmitting understanding than doing novel math, because they mostly exist in academia and that’s the job.


The futures I predict, out of my ass, for when the LLM wildfire spreads enough:

1. We are in a situation that the pace crumbles, in which situation we are left with classes of problems LLMs simply cannot touch because they have exhausted the "low hanging fruits" that their training and architecture was optimised for solving. In this scenario there is still stuff for the mathematicians to do, but the way young research mathematicians get trained has to be adjusted.

2. Or we simply see some slowdown, but it is still fast enough that it overtakes human capacities and produces novel mathematics that are harder and harder to understand, as LLMs are adapting to solve more and more types of problems faster than humans, while more and more fields are transcribed into lean. I guess if that happens the main role would be to clean up the llms' work, trying to get some gist out of it to build theories, and guide the llms for those with access to them. But not sure how many mathematicians will have access to these advanced LLMs to work with, so this may create big issues. Until now funding of mathematical work was somewhat trivial, they would only have expenses for traveling, conferences, and phds and postdocs, compared to other fields that need labs, equipment, high compute and/or research assistants to function.


When I was a kid in the 80s I noticed a lot of hackers of that era that typed like that. I thought it was strange at the time, but not at all uncommon


I've been doing something similar to this in a personal claude code frontend, though not particularly "magical".

I'm mostly using my system to make comments on long AI-generated documents (especially design documents). I find it works well to have the AI generate something, and then I read through it, making comments along the way.

You can get pretty far just repeating the things you see... "I'm reading [heading] and [comments]". But I do find some use in selecting content and saying "I don't agree with this" or whatever else.

The result is just an augmented message. It looks like:

    <transcript>
      Let's see what we've got here.
      <selection doc="proposal.md" location="paragraph 3">
        The system already...
      </selection>
      No, I don't like how this is approaching the problem, ...
    </transcript>
Then I just send this as a user message. Claude Code (and I'm guessing any of the agentic systems) picks up on the markup very easily. It also helps to label it as a transcript, as it can understand there may be errors, and things like spelling and punctuation are inferred not deliberate. (Some additional instruction is necessary to help it understand, for example, that it should look for homophones that might make more sense in context.)

It makes reviewing feel pretty relaxed and natural. I've played around with similar note taking systems, which I think could be great for studying in school, but haven't had the focus on that particular problem to take it very far.

But I think the best thing really is giving the agent a richer understanding of what the user is experiencing and doing and just creating a rich representation of that. The keywords can be useful, but almost only as checkpoints: a keyword can identify the moment to take the transcript and package it up and deliver it.

One difference perhaps in design motivation: I have really embraced long latency interactions. I use ChatGPT with extended thinking by default, and just suck it up when the answer didn't really require thinking. I deliver 10 points of feedback at once instead of little by little. (Often halfway through I explicitly contradict myself, because I'm thinking out loud and my ideas are developing.) I just don't stress out about latency or feedback, and so low-latency but lower-intelligence interactions don't do it for me (such as ChatGPT's advanced voice mode, or probably Thinking Machine's work). I think this focus is in part a value statement: I'm trying to do higher quality work, not faster work.


this is pretty built in with vscode plugins.

you select text in vscode, and write a comment, and the llm gets both


In a sense they do use their own language; they program in tokenized source, not ASCII source. And maybe that's just a form of syntactic sugar, like replacing >= with ≥ but x100. Or... maybe it's more than that? The tokenization and the models coevolve, from my understanding.

If we do enough passes of synthetic or goal-based training of source code generation, where the models are trained to successfully implement things instead of imitating success, then we may see new programming paradigms emerge that were not present in any training data. The "new language" would probably not be a programming language (because we train on generating source FOR a language, not giving it the freedom to generate languages), but could be new patterns within languages.


The knowledge machine question is fascinating ("Imagine you had access to a machine embodying all the collective knowledge of your ancestors. What would you ask it?") – it truly does not know about computers, has no concept of its own substrate. But a knowledge machine is still comprehensible to it.

It makes me think of the Book Of Ember, the possibility of chopping things out very deliberately. Maybe creating something that could wonder at its own existence, discovering well beyond what it could know. And then of course forgetting it immediately, which is also a well-worn trope in speculative fiction.


Jonathan Swift wrote about something we might consider a computer in the early 18th century, in Gulliver's Travels - https://en.wikipedia.org/wiki/The_Engine

The idea of knowledge machines was not necessarily common, but it was by no means unheard of by the mid 18th century, there were adding machines and other mechanical computation, even leaving aside our field's direct antecedents in Babbage and Lovelace.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: