Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

An endless loop of AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content and spitting out more content that also gets re-posted, resulting in a cesspool of BS masquerading as organic knowledge. I'm old enough to remember when Google provided meaningful search results rather than just SEO spam, the problem is about to get an order of magnitude worse.


IMO it's more vital than ever to fund projects like the Internet Archive. They're the only ones incentivized to maintain a snapshot of un-LLM-clouded training data of human knowledge, unclouded by the hubris of "who cares about the old stuff, we should focus our archiving on the web as it exists today" that inevitably will take hold (or already has) in big tech companies who will have laid off the vast majority of those voicing these concerns. We owe it to future generations to prevent ourselves from falling into the training-cycle trap.


The problem with the Internet Archive, which does an amazing job, is that they do an amazing job despite the problem being fundamentally intractable. Web content expands too quickly and too massively.

I wonder if the answer is a network of topic-focused archives; like moving from a "Library of Alexandria" model to a modern nationwide system of libraries.


” Web content expands too quickly and too massively.”

If most of it is crap I would call not archiving it a feature.

There is a weird convoluted analogue to CERN particle detectors. They smash particles together and then image the resulting storm of particle contrails via detector that is basically a sandwhiched ccd detector (like you have in camera, but different) the size of a cathedral. Resulting in far too much data for any system to analyze or even store in the first place. Hence they need/needed to runtime filter the massive amount of particle trail signals and only pick out the critical ones.

If there is too much data you simply need to drop the parts you are fairly confident you don’t need.

There is no reason there should be only one internet archive, there might very well be parallel operations filtering a bit different things.

I guess it’s a bit odd Unesco does not already have a parallel effort.


Okay so you build a knowledge graph on top of the internet archive. Now you are struggling to prioritize the resources necessary to capture long-tail content that doesn't mesh easily into popular corpuses. I imagine this would lead to the library equivalent of an echo chamber.


I was thinking more of a federated "webring" structure, with some content being present in more than one node, and where maintenance and curation are distributed (and gathered independently) among nodes.

The nation of, say, Japan, has limited interest in funding an american noprofit today; but they would likely have a great deal of interest in funding an equivalent focused on Japanese content, for example.


Ah so more like mastadon or ipfs, but specifically for the purposes of federated archiving.

So now you get into the issue of haves and have nots. Who is allowed to be considered an authorized archivist from a robots.txt perspective? Or what happens if an archivist becomes blacklisted for not respectfully crawling? How do national sanctions affect the Internet Archive of Russia? I imagine there would be a certification process and it would probably cost some money.

It's an interesting topic and I'm simply looking at the weak spots. I'm not against the overall concept though.


All legitimate questions, but if we only built perfect systems we would never have had TCP, let alone the pile of hacks we're now using to discuss this topic.

Distributed governance on the internet is a massive issue, and it's effectively unsolved for everything from pairing to DNS. In practice, good faith goes a long way, particularly in areas that are largely academic in scope - like archiving.


The curator being bandwidth-limited is not necessarily a problem if the problem you are solving is an overwhelmed audience in need of a curator. In other words, the Archive missing things may not really be a problem if the stuff is not missing is on average of value.

It raises the issue of governance of the curator, but the IA is already more transparent than Goole & co.


How do you measure the future value of something you don't keep?


would be nice if we could have a way to navigate just in the old web...


We still have curated libraries and the knowledge contained in those books has a much better signal-to-noise ratio than the internet ever will!


The Internet Archive is a library.


You're right, and it has a better signal to noise ratio than the internet in general, even when you factor in the Wayback Machine! Here's to curated knowledge!


The Wayback Machine is part of the library, too. Specifically, it's the periodicals room.


This! As data horder I see even more value in preserving pre-AI content, art, source code. That genuine stuff will be priceless.


You could call it: "AI Kessler syndrome".

The internet becomes so full of hallucinating AI output that it becomes impossible to train the model.


Worse, perhaps. Kessler Syndrome will eventually resolve itself as junk falls out of orbit over time, or new methods for cleaning it up are developed. Information, once buried in noise, becomes unrecoverable without a source of known truth for correlation.


Curation is the answer. The more junk, the better off are the one curation for quality, relevancy and human interest.


Curation that tracks the provenance. If we receive a string of text by itself we can't do much about it. We need to know from where it came from, whether it was written by a human, etc


Human provenance is less important than human curation here. AI can already infrequently output content that surpasses average human-generated quality in certain categories. As long as in the end you are checking that content exceeds an average bar of quality as assessed by human aesthetics, and ensuring that you have a diverse set of content (eg. not overrepresented by content that AI is particularly good or prolific at), it should still improve outcomes.


I still want to know if the content was written by an AI, because text quality (whatever that means) isn't the only metric humans care about.


This is a brilliant analogy


So the first one takes it all?


No more that as you get more AIs, all AIs get shittier together.


Or, the singularity.


That would be something else, like if someone built a Chat GPT that could train itself and it starts learning at an exponential rate, and learning how to make itself unstoppable by humans.


No AI will survive a solar flare! So we just need to hunker down in Zion and wait for that.


the singularity implies reaching a point where the AI’s improvement becomes self sustaining. this is the AI choking itself to death with bad training data.


I'm going to suggest giving Accelerando a read and consider the late state of the singularity.

http://www.accelerando.org

(heh - checking that link Charlie Stross just posted (Jan 31) a blog post: "An AI app walks into a writers room")

Anyways the link I was actually after ... give https://www.antipope.org/charlie/blog-static/fiction/acceler... a read and consider the "what happens in the later parts of the book."


In the back of my mind, I have a hope that it will lead to the collapse of the platform internet and a return to smaller trusted communities and boards.


That's part of the idea in The Expanding Dark Forest and Generative AI - https://maggieappleton.com/ai-dark-forest ( https://news.ycombinator.com/item?id=34243709 - 29 days ago; 419 comments)

> The dark forest theory of the web points to the increasingly life-like but life-less state of being online.Dark Forest Theory of the Internet by Yancey Strickler Most open and publicly available spaces on the web are overrun with bots, advertisers, trolls, data scrapers, clickbait, keyword-stuffing “content creators,” and algorithmically manipulated junk.

> It's like a dark forest that seems eerily devoid of human life – all the living creatures are hidden beneath the ground or up in trees. If they reveal themselves, they risk being attacked by automated predators.

> Humans who want to engage in informal, unoptimised, personal interactions have to hide in closed spaces like invite-only Slack channels, Discord groups, email newsletters, small-scale blogs, and digital gardens. Or make themselves illegible and algorithmically incoherent in public venues.


… or just talk to people in real life? (Until we have AI generated humans I suppose)


I see it making a similar progression as ads on radio and tv, homogenized mass media, paid product placements in shows. These AI generated content platforms are perfect for ads and social media propaganda. Mass customization.

I share the same hope but have doubts that as a society we'll have the collective critical thinking skills to disconnect from the AI overlords. We've already had the US and US inspired Brazilian coup attempts fueld by social media placements and it's only going to get more fine tuned and effective.

What can I do as an individual? One path is to simplify and declutter my digital life. How else to cope?


I hate being cynical, and I haven't really researched it, but my gut says there is too much invested from the non-tech world at this point in both the public and private sectors. The powers that be would probably rather force us all to asphyxiate on inane AI bullshit spewed out by our increasingly centrally-controlled technical world than cede control of electronic communication. Considering the pressure governments have freely applied to communications companies through policy, public shaming, disinformation, and secret infiltration (e.g. NSA breaking encryption to monitor gmail), and how effectively industry has skirted even the most basic privacy protections for users, I think they'll probably succeed. I don't see any reason to think that things like personal encryption or small community-run fora will change its course any more than legal guns have discouraged the creeping authoritarianism of the US Govt.


I don't have much knowledge around AI, but from what I can tell, it's dependent on inputs from across the web right? If so, then as the use of ChatGPT grows, it'll slowly get consumed back into the model. With enough iterations, it will start to veer more and more away from recognizable human speech/thought, like a recursive game of telephone.

Unless I'm completely misunderstanding how ML works, which very well may be true.


> Unless I'm completely misunderstanding how ML works, which very well may be true.

No, you got it right. Describing it as a game of telephone is a great analogy. This is exacerbated by the confidently incorrect problem. LLM output looks sophisticated and correct and may at time actually be correct. However some unpredictable percent of the time it will be incorrect and confidently so.


ChatGPT itself is being trained on curated content, it is clearly not trained on unscreened internet sites - this is easy enough to establish if you ask it questions around hot topic issues in the conspiracy groups - it gets the correct mainstream answer.

I suspect we will see the rise of both groups of machines, curated A.I.s and A.I.s just trained on anything, which should be entertaining.


Curated content by non native Enlish speakers earning as little as $2.00 an hour. What can possibly go wrong?


It gives you the correct mainstream answer by default. If you ask it to write, say, a hypothetical 4chan comment about such-and-such subject, and do sufficient prompt engineering to get past the filters, you'll see that it knows full well what the non-mainstream answers are:

https://i.imgur.com/u8Np332.png

The curation, such as it is, appears to be limited to humans downweighing the undesirable answers. Which is why there's always a way to work around it, even though it requires more and more elaborate prompts.


Unless it only reconsumes the “good” content. In other words, the stuff that got good reactions from humans. In which case, it will get better, not worse, at least at generating clickbait. But at least it will be coherent clickbait.


Not necessarily coherent. If the first generation model starts misusing a word or phrase or adopts a common misspelling, later generations of LLMs will pick up and amplify the error. Eventually you'll get purple Monkee dishwasher begging the question maps such as.


It will never be able to detect sarcasm amd satire and this will contribute greatly to the degradation of content quality , I think.

This will not work, its too soon.


Why do you think so? It's certainly capable of producing outputs along these lines if you tell it that this is what you want:

https://i.imgur.com/B2cHXRA.jpg

To me this implies that things such as "sarcasm" is a pattern simple enough for an AI to match - and that should go both ways, whether it's being generated or recognized.

If you're arguing that it won't be able to detect the more subtle sarcasm, then yeah, sure. But, well, Poe's Law predates GPT.


> AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content

Isn't that extrapolating the current trend a bit too much? Clearly, the text corpora[0] amassed before mass LLM content distribution are already big enough to train such models to decent general language fluency. So why would AI creators contaminate those datasets with potentially spurious content?

Sure, you want to keep your model up-to-date about the state of the world (the GPT corpus ends in mid-2021 afaik), but you can be much more careful about which texts you include. Those newer training data serve a different purpose than the original corpus, you don't need to bootstrap general language proficiency anymore. OpenAI already released a product for classifying AI-generated text, why would they not use something like that to filter future training data, for example?

[0] edited, thanks!


I mean we all know how adversarial training works as it's already commonly used today. I would not expect that to be useful in the future.


'corpora' fyi


Simulacra and Simulation comes to mind. Eventually, the real internet will cease to exist.


> An endless loop of AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content

Man I'm so tired of this very obvious observation. I wouldn't think a company smart enough to create an AI would also be dumb enough to fall into a pitfall that even the most casual observer can identify.


It may just be an arms race—the same AI that generates the nonsense also learns to identify and filter it out. Perhaps it'll also take down SEO spam with it. Feels unlikely but Google has an incentive to combat AI spam if people stop clicking on it, and presumably strong in-house AI capability…


In b4 controversial users get their content banned by accidentally tripping the AI filter


The other problematic area with that detection is people who have disabilities and use an AI assistant to help them type.


Excellent point


Absolutely, it's also going to get much more echo chambered and biased by corporate and political motivations.

And I'm not looking forward to a swathe of AI generated songs pumped into the charts and streaming services at potentially lots of songs per second.


Neil Stephenson predicted this in the novel Anathem. It ends up being an arms race with "bulshytt" filters versus generators.


Not quite the same, that scenario was deliberately engineered by spam filter vendors to make their filters necessary.

In the Kessler Syndrome analogy, that's an ablative aerospace impact armour company deliberately launching and blowing up satellites to sell their goods to spacecraft builders.


The worst part about this is that if there is another set of bots that tries to generate engagement, then the training data isn't coming from humans either. You have one set of actors spamming. And another set of actors upvoting stuff, predominantly their own but maybe also other random posts. So the resulting posts don't necessarily even cater to humans. It will be real online hellscape.


The web will transition strongly to verified identities, like we have with SSL certs. Along with filtering out people who use AI to post under a verified identity and get caught, It’s the only way to help ensure you’re reading actual human content.


Sorry, you can’t comment because your certificate was revoked when you died. You didn’t die?

——

We will pay you 100 bucks to withhold filing your partners death certificate for one week and providing their certificates to us.


That is how legally admissible e-signature schemes work. When the certificate holder dies or the certificate expires, it cannot any more be used to sign further documents.


Authoritarians are praying and dreaming for this to be realized.


If the web degenerates to the point where verified identities are required, then it really and truly will have died.


And with that and stylometric analysis, online pseudonymity effectively dies.


Both Google and Meta are massive targets.

Meta collapses if all its properties are filled with AI generated spam content.

Same for Google.

Meta will likely fall later since visual content at scale is still 12-18 months away. But for Google, the clock is ticking.


Why will Meta fall? Won’t people learn to unfollow accounts that post spam content? And if it is AI generated content that gets lots of engagement, wouldnt that help bring more engagement to Meta?


Agreed. At this rate isn't it just a matter of time before AI can easily get past CAPTCHAs? At which point someone will make the decision to start creating "organic" content advertising their product by getting an AI to write it at crazy speed across many different social media. Then the trick of appending "reddit.com" to Google search will die and we will begin to wonder if we are talking to real people on HN.


I am hoping it will be an improvement over the long form blog content people write to rank for SEO.

I don't want to read 10 pages of text for a simple cooking recipe.


Maybe information retrieval against unstructured text isn’t the right model? Maybe it’s time for google search to die.


AI generated content almost certainly will kill it as we know it. I don't expect the interface to change, but I expect Google's AI will "decide" what gets placed in search results, and where.


I think search itself was always a hack for how to ask human knowledge a question. It’s a great way to find a specific page of documentation. That’s really not what people want to do 99% or the time they use google.


Could tools that detect AI (like this one https://gptzero.substack.com/) be used as it consumes data and discard anything that gets flagged? I guess probably a cat and mouse game as more models get built though.


I mean that is already how AI works now. Adversarial models are mainstream.


And why will not some system arise where quality is valued. At my university, I hear a lot of colleges talk about how ChatGPT improves the quality of their work because they can find things they wouldn’t otherwise and because the writers block is partially solved.


This is most content on Youtube Shorts (reels?) now. In between Joe Rogan snippets are "Historical Photos" with voice-over descriptions and other rubbish. Very easy to imagine most of these being AI generated.

Its completely empty calories in terms of knowledge.


Sounds like the current web. What's the difference?


Better grammar and spelling.


One step closer to wireheads, with a trickle current going across a wire inserted into the pleasure center of the brain, in Ringworld.


I think the internet is already there with bots/content farms posting and upvoting garbage on social media.

Begun the Bot Wars have.


I suspect this will lead to a new wave of more advanced moderation techniques.


The torrent of garbage may require AI based moderation solution itself.

Although I hope we may see big come back of Web 1.0 forums such where users have to gain street cred and even invite referals with realy genuine contribution into community, no way to fake with AI today.


What is content? Is content entertainment? If you look at entertainment we will soon be living in world where short clips can be autogenerated and personalized based on keywords and criteria.

Stable Diffusion provides already good enough images to cover a lot of visual content online.

If you look at image shares sites for the purpose of entertainment like Imgur, you will also notice that a large portion of the viral content are screenshots from Twitter or traditional media.

Is content opinions? Most people don't have opinions on every topic on the planet. Is Gobekli Tepe the place of Noah's Ark? What's going on with Hunter Biden's laptop? How will Meta's VR strategy work out? Will it rain tomorrow in Sydney, Australia? Depending on your area of interest you might or might not have an opinion about it which you may or may not publish online.


ChatGpt generated comment,?


That's what ChatGPT had to say about this : Content refers to information, experiences, or resources that are created to serve a specific audience and purpose. This can include text, images, audio, video, or other forms of media. The content can be used for a variety of purposes, such as education, entertainment, marketing, or news.


Infinitely-expanding snake eats infinitely-expanding tail.


One solution is pretty simple: pedigree.

Divide up the 'net into trusted and untrusted sources. Make the trust ratings public. Use search tools and corpuses such as the Google Books dataset to source "knowledge" back to pre-Internet roots, when necessary. In short: bring academic reputation back and bring it back hard.

It will make for a more elitist web, but given that even without ChatGPT we've had a problem with wildfire misinformation spread in social media networks it might be a change that's a long time coming.


That means anonymously posting on stackoverflow will be gone...


Stackoverflow already has a pretty solid reputation system with pseudonymous users.

What it means is the bar for becoming a new StackOverflow contributor (or Reddit admin, or Wikipedian) might become much, much higher. "Oh, you want to contribute your first post? Show me the bicycles in this image, find the letters in this image, and provide the names of two existing Stack Overflow users with over 1000 karma who can vouch for you, and also you see a tortoise on its back, baking in the sun. You're not helping it. Why aren't you helping it?..."




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: