An endless loop of AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content and spitting out more content that also gets re-posted, resulting in a cesspool of BS masquerading as organic knowledge. I'm old enough to remember when Google provided meaningful search results rather than just SEO spam, the problem is about to get an order of magnitude worse.
IMO it's more vital than ever to fund projects like the Internet Archive. They're the only ones incentivized to maintain a snapshot of un-LLM-clouded training data of human knowledge, unclouded by the hubris of "who cares about the old stuff, we should focus our archiving on the web as it exists today" that inevitably will take hold (or already has) in big tech companies who will have laid off the vast majority of those voicing these concerns. We owe it to future generations to prevent ourselves from falling into the training-cycle trap.
The problem with the Internet Archive, which does an amazing job, is that they do an amazing job despite the problem being fundamentally intractable. Web content expands too quickly and too massively.
I wonder if the answer is a network of topic-focused archives; like moving from a "Library of Alexandria" model to a modern nationwide system of libraries.
” Web content expands too quickly and too massively.”
If most of it is crap I would call not archiving it a feature.
There is a weird convoluted analogue to CERN particle detectors. They smash particles together and then image the resulting storm of particle contrails via detector that is basically a sandwhiched ccd detector (like you have in camera, but different) the size of a cathedral. Resulting in far too much data for any system to analyze or even store in the first place. Hence they need/needed to runtime filter the massive amount of particle trail signals and only pick out the critical ones.
If there is too much data you simply need to drop the parts you are fairly confident you don’t need.
There is no reason there should be only one internet archive, there might very well be parallel operations filtering a bit different things.
I guess it’s a bit odd Unesco does not already have a parallel effort.
Okay so you build a knowledge graph on top of the internet archive. Now you are struggling to prioritize the resources necessary to capture long-tail content that doesn't mesh easily into popular corpuses. I imagine this would lead to the library equivalent of an echo chamber.
I was thinking more of a federated "webring" structure, with some content being present in more than one node, and where maintenance and curation are distributed (and gathered independently) among nodes.
The nation of, say, Japan, has limited interest in funding an american noprofit today; but they would likely have a great deal of interest in funding an equivalent focused on Japanese content, for example.
Ah so more like mastadon or ipfs, but specifically for the purposes of federated archiving.
So now you get into the issue of haves and have nots. Who is allowed to be considered an authorized archivist from a robots.txt perspective? Or what happens if an archivist becomes blacklisted for not respectfully crawling? How do national sanctions affect the Internet Archive of Russia? I imagine there would be a certification process and it would probably cost some money.
It's an interesting topic and I'm simply looking at the weak spots. I'm not against the overall concept though.
All legitimate questions, but if we only built perfect systems we would never have had TCP, let alone the pile of hacks we're now using to discuss this topic.
Distributed governance on the internet is a massive issue, and it's effectively unsolved for everything from pairing to DNS. In practice, good faith goes a long way, particularly in areas that are largely academic in scope - like archiving.
The curator being bandwidth-limited is not necessarily a problem if the problem you are solving is an overwhelmed audience in need of a curator. In other words, the Archive missing things may not really be a problem if the stuff is not missing is on average of value.
It raises the issue of governance of the curator, but the IA is already more transparent than Goole & co.
You're right, and it has a better signal to noise ratio than the internet in general, even when you factor in the Wayback Machine! Here's to curated knowledge!
Worse, perhaps. Kessler Syndrome will eventually resolve itself as junk falls out of orbit over time, or new methods for cleaning it up are developed. Information, once buried in noise, becomes unrecoverable without a source of known truth for correlation.
Curation that tracks the provenance. If we receive a string of text by itself we can't do much about it. We need to know from where it came from, whether it was written by a human, etc
Human provenance is less important than human curation here. AI can already infrequently output content that surpasses average human-generated quality in certain categories. As long as in the end you are checking that content exceeds an average bar of quality as assessed by human aesthetics, and ensuring that you have a diverse set of content (eg. not overrepresented by content that AI is particularly good or prolific at), it should still improve outcomes.
That would be something else, like if someone built a Chat GPT that could train itself and it starts learning at an exponential rate, and learning how to make itself unstoppable by humans.
the singularity implies reaching a point where the AI’s improvement becomes self sustaining. this is the AI choking itself to death with bad training data.
In the back of my mind, I have a hope that it will lead to the collapse of the platform internet and a return to smaller trusted communities and boards.
> The dark forest theory of the web points to the increasingly life-like but life-less state of being online.Dark Forest Theory of the Internet by Yancey Strickler Most open and publicly available spaces on the web are overrun with bots, advertisers, trolls, data scrapers, clickbait, keyword-stuffing “content creators,” and algorithmically manipulated junk.
> It's like a dark forest that seems eerily devoid of human life – all the living creatures are hidden beneath the ground or up in trees. If they reveal themselves, they risk being attacked by automated predators.
> Humans who want to engage in informal, unoptimised, personal interactions have to hide in closed spaces like invite-only Slack channels, Discord groups, email newsletters, small-scale blogs, and digital gardens. Or make themselves illegible and algorithmically incoherent in public venues.
I see it making a similar progression as ads on radio and tv, homogenized mass media, paid product placements in shows. These AI generated content platforms are perfect for ads and social media propaganda. Mass customization.
I share the same hope but have doubts that as a society we'll have the collective critical thinking skills to disconnect from the AI overlords. We've already had the US and US inspired Brazilian coup attempts fueld by social media placements and it's only going to get more fine tuned and effective.
What can I do as an individual?
One path is to simplify and declutter my digital life. How else to cope?
I hate being cynical, and I haven't really researched it, but my gut says there is too much invested from the non-tech world at this point in both the public and private sectors. The powers that be would probably rather force us all to asphyxiate on inane AI bullshit spewed out by our increasingly centrally-controlled technical world than cede control of electronic communication. Considering the pressure governments have freely applied to communications companies through policy, public shaming, disinformation, and secret infiltration (e.g. NSA breaking encryption to monitor gmail), and how effectively industry has skirted even the most basic privacy protections for users, I think they'll probably succeed. I don't see any reason to think that things like personal encryption or small community-run fora will change its course any more than legal guns have discouraged the creeping authoritarianism of the US Govt.
I don't have much knowledge around AI, but from what I can tell, it's dependent on inputs from across the web right? If so, then as the use of ChatGPT grows, it'll slowly get consumed back into the model. With enough iterations, it will start to veer more and more away from recognizable human speech/thought, like a recursive game of telephone.
Unless I'm completely misunderstanding how ML works, which very well may be true.
> Unless I'm completely misunderstanding how ML works, which very well may be true.
No, you got it right. Describing it as a game of telephone is a great analogy. This is exacerbated by the confidently incorrect problem. LLM output looks sophisticated and correct and may at time actually be correct. However some unpredictable percent of the time it will be incorrect and confidently so.
ChatGPT itself is being trained on curated content, it is clearly not trained on unscreened internet sites - this is easy enough to establish if you ask it questions around hot topic issues in the conspiracy groups - it gets the correct mainstream answer.
I suspect we will see the rise of both groups of machines, curated A.I.s and A.I.s just trained on anything, which should be entertaining.
It gives you the correct mainstream answer by default. If you ask it to write, say, a hypothetical 4chan comment about such-and-such subject, and do sufficient prompt engineering to get past the filters, you'll see that it knows full well what the non-mainstream answers are:
The curation, such as it is, appears to be limited to humans downweighing the undesirable answers. Which is why there's always a way to work around it, even though it requires more and more elaborate prompts.
Unless it only reconsumes the “good” content. In other words, the stuff that got good reactions from humans. In which case, it will get better, not worse, at least at generating clickbait. But at least it will be coherent clickbait.
Not necessarily coherent. If the first generation model starts misusing a word or phrase or adopts a common misspelling, later generations of LLMs will pick up and amplify the error. Eventually you'll get purple Monkee dishwasher begging the question maps such as.
To me this implies that things such as "sarcasm" is a pattern simple enough for an AI to match - and that should go both ways, whether it's being generated or recognized.
If you're arguing that it won't be able to detect the more subtle sarcasm, then yeah, sure. But, well, Poe's Law predates GPT.
> AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content
Isn't that extrapolating the current trend a bit too much? Clearly, the text corpora[0] amassed before mass LLM content distribution are already big enough to train such models to decent general language fluency. So why would AI creators contaminate those datasets with potentially spurious content?
Sure, you want to keep your model up-to-date about the state of the world (the GPT corpus ends in mid-2021 afaik), but you can be much more careful about which texts you include. Those newer training data serve a different purpose than the original corpus, you don't need to bootstrap general language proficiency anymore. OpenAI already released a product for classifying AI-generated text, why would they not use something like that to filter future training data, for example?
> An endless loop of AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content
Man I'm so tired of this very obvious observation. I wouldn't think a company smart enough to create an AI would also be dumb enough to fall into a pitfall that even the most casual observer can identify.
It may just be an arms race—the same AI that generates the nonsense also learns to identify and filter it out. Perhaps it'll also take down SEO spam with it. Feels unlikely but Google has an incentive to combat AI spam if people stop clicking on it, and presumably strong in-house AI capability…
Not quite the same, that scenario was deliberately engineered by spam filter vendors to make their filters necessary.
In the Kessler Syndrome analogy, that's an ablative aerospace impact armour company deliberately launching and blowing up satellites to sell their goods to spacecraft builders.
The worst part about this is that if there is another set of bots that tries to generate engagement, then the training data isn't coming from humans either. You have one set of actors spamming. And another set of actors upvoting stuff, predominantly their own but maybe also other random posts. So the resulting posts don't necessarily even cater to humans. It will be real online hellscape.
The web will transition strongly to verified identities, like we have with SSL certs. Along with filtering out people who use AI to post under a verified identity and get caught, It’s the only way to help ensure you’re reading actual human content.
That is how legally admissible e-signature schemes work. When the certificate holder dies or the certificate expires, it cannot any more be used to sign further documents.
Why will Meta fall? Won’t people learn to unfollow accounts that post spam content? And if it is AI generated content that gets lots of engagement, wouldnt that help bring more engagement to Meta?
Agreed. At this rate isn't it just a matter of time before AI can easily get past CAPTCHAs? At which point someone will make the decision to start creating "organic" content advertising their product by getting an AI to write it at crazy speed across many different social media. Then the trick of appending "reddit.com" to Google search will die and we will begin to wonder if we are talking to real people on HN.
AI generated content almost certainly will kill it as we know it. I don't expect the interface to change, but I expect Google's AI will "decide" what gets placed in search results, and where.
I think search itself was always a hack for how to ask human knowledge a question. It’s a great way to find a specific page of documentation. That’s really not what people want to do 99% or the time they use google.
Could tools that detect AI (like this one https://gptzero.substack.com/) be used as it consumes data and discard anything that gets flagged? I guess probably a cat and mouse game as more models get built though.
And why will not some system arise where quality is valued. At my university, I hear a lot of colleges talk about how ChatGPT improves the quality of their work because they can find things they wouldn’t otherwise and because the writers block is partially solved.
This is most content on Youtube Shorts (reels?) now. In between Joe Rogan snippets are "Historical Photos" with voice-over descriptions and other rubbish. Very easy to imagine most of these being AI generated.
Its completely empty calories in terms of knowledge.
The torrent of garbage may require AI based moderation solution itself.
Although I hope we may see big come back of Web 1.0 forums such where users have to gain street cred and even invite referals with realy genuine contribution into community, no way to fake with AI today.
What is content?
Is content entertainment?
If you look at entertainment we will soon be living in world where short clips can be autogenerated and personalized based on keywords and criteria.
Stable Diffusion provides already good enough images to cover a lot of visual content online.
If you look at image shares sites for the purpose of entertainment like Imgur, you will also notice that a large portion of the viral content are screenshots from Twitter or traditional media.
Is content opinions?
Most people don't have opinions on every topic on the planet.
Is Gobekli Tepe the place of Noah's Ark?
What's going on with Hunter Biden's laptop?
How will Meta's VR strategy work out?
Will it rain tomorrow in Sydney, Australia?
Depending on your area of interest you might or might not have an opinion about it which you may or may not publish online.
That's what ChatGPT had to say about this :
Content refers to information, experiences, or resources that are created to serve a specific audience and purpose. This can include text, images, audio, video, or other forms of media. The content can be used for a variety of purposes, such as education, entertainment, marketing, or news.
Divide up the 'net into trusted and untrusted sources. Make the trust ratings public. Use search tools and corpuses such as the Google Books dataset to source "knowledge" back to pre-Internet roots, when necessary. In short: bring academic reputation back and bring it back hard.
It will make for a more elitist web, but given that even without ChatGPT we've had a problem with wildfire misinformation spread in social media networks it might be a change that's a long time coming.
Stackoverflow already has a pretty solid reputation system with pseudonymous users.
What it means is the bar for becoming a new StackOverflow contributor (or Reddit admin, or Wikipedian) might become much, much higher. "Oh, you want to contribute your first post? Show me the bicycles in this image, find the letters in this image, and provide the names of two existing Stack Overflow users with over 1000 karma who can vouch for you, and also you see a tortoise on its back, baking in the sun. You're not helping it. Why aren't you helping it?..."