Search

Items tagged with: llms


"GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels."

They go on to say "GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness", and "A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions."

If you're unfamiliar with the ARC-AGI tests, they're pattern-matching tests that are intended to be easy for humans but hard for AI despite AI being able to solve math problems, etc, that are hard for humans. It's intended to give a more realistic idea how close AI models are to surpassing human intelligence at more or less everything, not just specific tasks that are hard for humans. This threshold is called "artificial general intelligence" (AGI) which is not exactly an intuitive term (but math and science and the AI field are full of terms that are not intuitive -- can you come up with a better one?). ARC-AGI-3 is the 3rd version of the test, because AI keeps getting better and they keep having to make the test harder.

"These environments only contain core knowledge priors and are difficulty-calibrated through controlled testing with human participants. Humans can solve 100% of the environments."

"The goal of the ARC-AGI series is to measure the 'residual gap' between current artificial intelligence and AGI. We define AGI as a system's ability to acquire any skill a human can, as efficiently as a human can."

If you're wondering what the explanation is for the dollar amounts in the description above, they say: "For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses."

"Most of this fee pays for the participant's time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain's energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

"Beyond the scores, Astra's replays show how it turns unfamiliar game mechanics into useful working models. Three findings stood out: the compact algebraic notation it develops, its action efficiency compared with humans, and the custom tools it builds."

If you're wondering what the bit about "action efficiency" is all about, they say: "For each level, we defined the 'human baseline' using the median action count among players who completed it. This gives us a reference for comparing human and AI performance. An AI that needs more actions is less action-efficient, while one that needs fewer actions is more action-efficient."

"Before we launched ARC-AGI-3, we hypothesized that action efficiency would remain a dividing line between humans and AI. We anticipated that even when an AI solved an environment, it might require substantially more exploration (actions) than a person. That remains true of brute-force approaches, but frontier AI shows a more binary-like pattern. Once frontier AI 'understands' the mechanics, it generally executes within the range of human efficiency."

"Astra's results show that it needed fewer interactions than the human baseline to execute a solution."

There's also some stuff about "harnesses". They say: "The Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment", and "The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work."

OpenAI's GPT-6 Astra on ARC-AGI-3

#solidstatelife #ai #genai #llms #agi


jev-lint (no capitalization) is a static analyzer-ish program that analyzes source code -- only JavaScript and TypeScript -- against a document in English of conventions the code has to follow, and generates error messages that are intended for your code-generating AI agent to read. As the name implies, it uses Jev.

zdenham / jev-lint

#solidstatelife #ai #genai #llms #codingai #rlcd #jev #staticanalysis



The media in this post is not displayed to visitors. To view it, please go to the original post.

‘Doom Loop’: OpenAI and Microsoft Admits #LLMs Are Destroying the Web and Built on Theft

"Executives working on #AI at #Microsoft and #OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing #theft of unprecedented proportions,” and the “largest theft of #labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire #web.”

404media.co/doom-loop-openai-a…


‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft


Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”

Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.

It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.

“Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions,” an internal Microsoft document cited in the case read, adding “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”

The filing was written by lawyers for the New York Times but is largely comprised of statements and interviews with big tech executives that admit both that LLMs are largely trained on stolen content, that they represent an existential risk for the human writers, artists, and media companies that they stole from, and that their products have started a “doom loop” that is eating the web and destroying the businesses that these companies stole from.

The unredacted filing was found by Jason Kint, the CEO of Digital Content Next, a trade organization that represents digital media companies. Kint has been closely following and posting about massive AI copyright lawsuits.

The court filing cites an internal Microsoft document that found that AI products steal from human content creators, then cannibalize clicks from the people and websites they’ve stolen from, thereby destroying their business models.

“Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said.

Microsoft executives, including CEO Satya Nadella, testified under oath that after ripping content from the New York Times and other news sites, clicks to those news sites fully cratered, falling by more than 90 percent on Bing.

Documents obtained during the court proceedings found that OpenAI created “a hack to get around nytimes paywall,” to which OpenAI cofounder Greg Brockman said “ah, nice.” Microsoft executive Brent Hecht wrote that LLMs steal content “without ways of distributing economic value down the supply chain, [which] necessarily threatens the economic stability of those who create the content.”

OpenAI’s policy director Jack Clark wrote that the company was “creating systems that substitute for the labor of the people that define the ‘culture’ of society” and Microsoft, in a policy document, wrote that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained […] LLMs are a product that destroys its supply chain.” OpenAI called itself an “existential threat” to news publishers, and an OpenAI software engineer testified that “no matter how prominently we show the links, users won’t click.”

None of this is at all surprising to anyone who has been paying attention to the development of generative artificial intelligence, but the document, taken in whole, is a real they-admit-it situation. OpenAI’s and Microsoft’s lawyers have been trying to argue that their model training is fair use and transformative under copyright law and that they are building something that is fundamentally different from the human labor that it was trained on. But internally, these executives know that what they have built has been built on stolen content and that the products they’ve made are cannibalizing the sources they’ve stolen from and destroying the internet as we know it.



Maybe AI is going to kill is all… via a meat proxy. Basically, having an LLM anywhere near decisions that might lead to military strikes is plain stupidity coupled with a stunning absence of ethics.

In the 80s, Stanislav Petrov went against protocol and prevented a nuclear launch that no doubt would’ve ended with millions of people dead.

Today? We’re in real danger of the same scenario going the other way because people are placing unfounded and dangerous amounts of trust in systems that are unproven, inaccurate, and designed to lie convincingly while trying to ingratiate themselves with their users in order to encourage more usage.

From the story: According to the report, "The US military swung into action with plans to intercept the vessel, according to four sources familiar with the episode. According to two of the sources, armed members of the US military were preparing to board the ship. Military planes were in the air, one of those sources and another source familiar with the incident said. It was only just before the planned operation that officials dug deeper into the report put together by a special operations command analyst and found it had been generated with the help of artificial intelligence (AI)."

#LLMs #AI #WarGames #OhGodOhGodWeAreAllGonnaDie
rawstory.com/penatgon-ai-war-c…


Eine bekannte Person hat von deren Therapeutin einen Zettel vorgelegt bekommen, dass Sitzungen und Zukunft von einer #KI der Firma #VIAHealthTech mitgeschnitten und ausgewertet werden, sie soll das unterschreiben.

Laut ViaHeathTech alles DSGVO konform und Schweigepflicht wird gewahrt.

Weiß er was dazu finde wenig Infos die sich kritisch damit auseinander setzen.

#fedihelp #ai #llms #Datenschutz #medizin #psychologie #therapie #psychotherapie


"I'm not sure exactly when it happened. As recently as December 2025, I still found some Claude-isms kind of catchy and clever -- I noticed AI language, but it didn't trigger violent rage. By spring, I was snapping at the tendons of any poor soul who showed up in my inbox with an 'I'd value your take on this' or 'the call most leaders still won't make'."

Wait, AI says "I'd value your take on this" and "the call most leaders still won't make"?

Let's continue.

"It started with DMs and emails, but it didn't stop there. By summer, my reactive rage-response to AI-generated text had spread to include most forms of writing. If I'm reading a newsletter and I start to sense AI-isms, I delete and unsubscribe. If I'm reading a blog post, I close the tab; if I'm on social media, I unfollow or unfriend. If it happens repeatedly, I will go out of my way to avoid that writer in the future. I mostly try to not engage, but if I had a button that would let me deliver a 10,000 volt shock to the author I would slam that button every time and I wouldn't care who saw."

While I can relate to the feeling of realizing something is probably AI generating and wanting to stop reading, there is the problem of AI being so good at imitating humans, it's genuinely hard to tell. There's actually research on this. I should track some of those studies down. People have been given text written by AI and humans and challenged to tell which is written by AI and which is written by humans. The same thing has been done with art: tell which art image is made by humans and which is made by AI. Humans can't tell the difference. The Turing Test has been passed.

"Sometimes, yes, the quality of the work is all that matters."

"Other times, the value of a piece of writing derives from the fact that a particular person said it, thought it or felt it, or its value is grounded in your relationship."

She diagrams out a "personal" to "functional" continuum.

Confessions of an unrepentant slop snob

#solidstatelife #ai #genai #llms #aislop


my first mention of this on Mastodon goes back to shortly when i joined in 2023.

mastodon.social/@blogdiva/1106…

but been saying this for decades ―all anthropomorphic robotics is a reinvention of #slavery, and that includes #ai #LLMs

here’s more of what i’ve posted onhere:

mastodon.social/@blogdiva/1106…
mastodon.social/@blogdiva/1111…
mastodon.social/@blogdiva/1112…
mastodon.social/@blogdiva/1144…

@SallyStrange


OpenArch is "PyTorch implementations of modern open-source LLM architectures (Llama, Qwen, DeepSeek, Gemma, GPT-OSS, Kimi, and more) -- written from scratch for readability and learning, based on Sebastian Raschka's LLM Architecture Gallery."

Looks like I super valuable resource for all of you wanting to learn LLMs. (I do too but I don't know where to find the time.)

"This repository contains hand-written PyTorch implementations of the model architectures cataloged in Sebastian Raschka's LLM Architecture Gallery. Each model is implemented to the best of my knowledge from the original papers, technical reports, reference config.json files, and the excellent writeups by Sebastian Raschka and Machine Learning Mastery."

"The goal is not to compete with transformers or other production libraries. The goal is clarity and learning: a single readable file per architecture, with the structural choices (attention type, normalization, layer mix, MoE routing, positional encoding) made explicit and easy to compare side-by-side."

"Modern LLM architectures share a common skeleton but differ in dozens of small, important choices:"

  • "Attention: MHA, GQA, MQA, MLA, sliding-window, linear/DeltaNet hybrids"
  • "Normalization: pre-norm, post-norm, QK-Norm, sandwich norm, RMSNorm"
  • "Positional encodings: RoPE, NoPE, partial RoPE, YaRN"
  • "Decoder type: dense vs sparse MoE (with or without shared experts), hybrid Mamba/attention"
  • "Training-time tricks: Multi-token-prediction, latent experts, gated attention"

"Reading the official model code can be hard because production repos optimize for speed, sharding, and backward compatibility. This repo optimizes for reading."

anuj0456 / OpenArch

#solidstatelife #ai #genai #llms #pytorch


The Segway is here to stay, you know! Once someone invents something, the cat's out of the bag! Oh you don't use a Segway? Watch out, or this amazing new technology will leave you (slightly) behind! You'd better invest now you know! This technology is poised to revolutionize personal transportation! If you don't get in early, you'll never catch up and make the big bucks (unless you can manage a brisk jog or have a car or something). I don't know what they're teaching in PE these days, but they definitely need to start teaching how to ride a Segway! It's completely intuitive, you see! Should athletes be allowed to ride Segways around the track instead of running laps? Of course! After all, sports are already starting to shift to Segway leagues. These things *are* the future! They're totally safe,* even around cliffs!

*Remember to always wear your helmet when riding, and be prepared to land on your feet like a cat if user error results in an unexpected temporary levitation event. Segway, inc. is not liable for damages caused by failure to deploy superhuman reflexes at a moment's notice.

This post contains some hyperbole, but not a much as you might think. For the curious: en.wikipedia.org/wiki/Segway#H…

#AI #LLMs #GenAI


1/ Dan Flickinger pointed me to this paper about LLMs. The authors used #LLMs to generate NYT style text and compared the results with texts written by humans. They parsed the samples with the #ERG, a large-scale #HPSG grammar for #English, written by Dan. They could show that the text by humans is more diverse, that certain low frequency constructions do not appear in LLM-generated texts and that the LLMs are more similar to each other than to the individual authors.

#AI



@ambiguous_yelp They proudly tooted about starting LLM reviews the day before yesterday. That caused an outrage because why use GrapheneOS at all if you don't care about security?! Can't point you to it because I'm blocked. Can anyone else help?

#LLM #LLMs #AI #noAI #technofascism #grapheneos


Bullshit. Sowohl #Anthropic als auch #OpenAI schüren das Märchen von der bevorstehenden Apokalypse und dass sie noch mehr Zeit und damit Geld benötigen, um alles schön kontrolliert zu entwickeln. Die angeblich bestehende Gefahr wurde nie konkretisiert, #LLMs besitzen kein Bewusstsein und den Tech-Bros fehlt es an Alternativen zu #Chatbots & Co., wohin Aufmerksamkeit und Geld gehen könnten. Also wird der zähe KI-Gummi weiter gezogen, noch mehr Investoren geworben und die bestehenden vertröstet, damit der schöne Kreislauf des Geldglücksrads nicht unterbrochen wird und das Kartenhaus (noch) nicht zusammenfällt. Wir sehen einem todkranken Patienten dabei zu, wie er nur weiter beatmet wird.


A solution to a longstanding math problem relating to the Navier-Stokes equations has been found by AI. The Navier-Stokes equations are equations for taking Newton's laws of motion, which work for discrete objects, and making them continuous, so they apply to fluids, and making them continuously feed back on themselves, so they take the form of differential equations. I can't explain it better than the summary from OpenAI's blog post, so I'll just quote the blog post.

"The equations date to the nineteenth-century work of Claude-Louis Navier and George Gabriel Stokes. In 1934, Jean Leray proved that solutions exist in a generalized sense, but whether they always remain smooth became a central unanswered question. In 2000, the Clay Mathematics Institute named the Navier-Stokes existence and smoothness problem one of seven Millennium Prize Problems."

"The Navier-Stokes equations use Newton's second law of motion ('F=ma') to describe how fluids move. Importantly, they treat a fluid as a continuous medium rather than tracking individual molecules. These equations are used for aircraft design, weather forecasting, and the study of blood flow."

"A fundamental open question for these dynamical equations has been whether the continuum approximation of the fluid can break down. Specifically, can the Navier-Stokes equations for a three-dimensional incompressible fluid with constant density develop a 'singularity,' even when the motion starts smoothly? Here, a singularity means the dynamics lead to speeds in the fluid growing without bound within a finite amount of time. The development of a singularity would have to happen despite the presence of viscosity, which tends to smooth out motion. Because a real fluid cannot move infinitely fast, this would mark a breakdown in how the equations model the fluid. To continue modeling the system, one would then need to track the behaviour of each particle individually."

The AI found a singularity, and came up with a formalized proof in the form of a Lean proof written for the proof assistant software Lean, which has verified the proof as correct.

"The solution is a vortex, a spinning swirl of fluid, that spirals inward and gets increasingly elongated, like spaghetti. This central region shrinks while it speeds up in such a way that its energy still stays finite, as required by the laws of physics. The technical challenge is for the equations to develop the breakdown through the motion of the fluid itself, rather than, for example, us putting in an infinite force by hand. More mathematically, the terms in the Navier-Stokes equations that describe the motion -- acceleration, pressure gradients, momentum transfer, viscosity -- must both become big yet cancel in a precise way. This detailed balance leaves a smooth external force even as the velocity of the fluid grows without bound."

Interestly, though the Clay Mathematics Institute offers $1 million for the solution to this problem, because the solution was found by AI and not a human (or group of humans), nobody will receive the $1 million prize. AIs, as it turns out, are not eligible for the prize money. I wonder if this means the age of humans getting prize money for solving math problems is drawing to a close, since going forward, it will be impossible to know if a math problem was solved entirely by humans, entirely by AI, or anywhere in between.

This problem was not solved by a single AI agent, either. It was solved by a swarm of AI agents, which were all instances of a new model more powerful than GPT-6 Astra which has recently been made available to the public. OpenAI doesn't say how many AI agents were in the swarm, or how exactly they were coordinated by other AI agents.

There's another aspect to this as well. A pair of mathematicians are wondering if their conversations with ChatGPT were used by the model to help OpenAI arrive at this proof. OpenAI says there was nothing beyond the ordinary user interaction feedback that is used to help train the next generation of models. Interestingly it looks like OpenAI can't prove the conversations with those mathematicians didn't contribute to this proof at all.

On the Navier–Stokes Millennium Prize Problem

#solidstatelife #ai #genai #llms #codingai #mathematics #lean


"Coding is expansive while other white-collar work is compressive. A short program, or these days a short prompt, can and does trigger planning, code generation, testing, debugging, retries and repeated ingestion of a large codebase followed by a great deal of computation and reams and reams of output. In a world in which the computer will try twenty versions of a command before it hits upon the correct format, token creation and use explodes as the machine groups toward an answer. By contrast, the business of a lawyer is to take the expansive set of legal codes and case situations that is the law and squeeze it down into a brief, an opinion, a recommendation. The business of a consultant is much the same. And the whole point of management is to throw away as much information as you can in order to make the problems of direction and coordination graspable and actionable. Summarization, document review, research synthesis, meeting notes and similar tasks. Large inputs, and relatively small outputs. No explosion of agentic activity once the universe of input documents has been defined and collected."

Brad DeLong reacts to Paul Kedrosky's claim that software developers are highly unrepresentative of broader AI use.

"As Paul Kedrosky says, AI could become ubiquitous across law, finance, and consulting, yet those workers would still burn far fewer tokens each than software developers do."

What about art generation? Video generation? Text-to-voice? Music generation? What about writers? What about all the boring writing like those long legal contracts that I'm sure none of you ever click through without reading (right?) or technical documentation? And what about once robotics really gets going, and every "action" burns whatever the future equivalent of "action tokens" will be? (Ok, ok, he said "white collar".)

How special is programming as a token-infall attractor?: Chart of the Day

#solidstatelife #ai #genai #llms #codingai


The media in this post is not displayed to visitors. To view it, please go to the original post.

US government argues that fair use applies to OpenAI’s training of LLMs in key copyright case
Historically, legislation and court decisions have always lagged behind technological change. But the gap between the two is even starker in the world of generative AI. The large language model (LLM) that kicked things off, ChatGPT, was only released in November 2022; the pace of development since then has been dizzying. Meanwhile, nearly 200 lawsuits have been brought against AI companies, […]
#andresGuadamuz #ani #chatgpt #competition #delhi #eff #eu #fairUse #genai #india #injunction #judge #law #lawsuits #licensing #llms #marketDilution #nationalSecurity #newYorkTimes #oligopoly #openai #privateUse #technollama #usConstitution #usDepartmentOfJusticewalledculture.org/us-governmen…


The media in this post is not displayed to visitors. To view it, please go to the original post.

US government argues that fair use applies to #OpenAI’s training of #LLMs in key #copyright case - walledculture.org/us-governmen…


US government argues that fair use applies to OpenAI’s training of LLMs in key copyright case
Historically, legislation and court decisions have always lagged behind technological change. But the gap between the two is even starker in the world of generative AI. The large language model (LLM) that kicked things off, ChatGPT, was only released in November 2022; the pace of development since then has been dizzying. Meanwhile, nearly 200 lawsuits have been brought against AI companies, […]
#andresGuadamuz #ani #chatgpt #competition #delhi #eff #eu #fairUse #genai #india #injunction #judge #law #lawsuits #licensing #llms #marketDilution #nationalSecurity #newYorkTimes #oligopoly #openai #privateUse #technollama #usConstitution #usDepartmentOfJusticewalledculture.org/us-governmen…



Thomas Kwa left METR to take a job at OpenAI, but that's not the interesting part of this bit, the interesting part is that he wrote about it and said the reason he made the switch and what he'll be doing on his new job is "measuring and modeling recursive self-improvement".

"Why did I leave METR? Briefly, I want to inform the world whether recursive self-improvement (RSI) is imminent, which requires modeling RSI, which I think I can do better at OpenAI."

"I think OpenAI currently lacks the strategic awareness and thoughtfulness needed to responsibly build a technology with the extreme downside risk of artificial superintelligence (ASI). If they somehow cause a singularity in 6 months without applying any lessons learned from the HuggingFace incident, misaligned takeover would seem more likely than not."

"We probably don't have self-sustaining acceleration now, but we will probably get 3+ years of progress in 8 months at some point when AIs can fully substitute for humans -- and 3+ years of progress in 3 months is plausible due to superhuman inference scaling."

He says nothing about what will actually be measured to "measure recursive self-improvement", but see below for an update from OpenAI themselves that reveals the things they have started measuring.

Thomas Kwa's Shortform -- LessWrong

#solidstatelife #ai #genai #llms #codingai #agi #recursiveselfimprovement #rsi


hachyderm.io/@eliasulrich/1172… eliasulrich@hachyderm.io - #Harvard published a paper with a devastating title: “Large-Language Models as a #CognitiveVirus”

"Think about how a #virus operates:

It cannot replicate on its own. It requires a host cellular machinery to copy itself.

An #LLM cannot execute, compute, or spread on its own. It requires human cognition, human servers, and human networks to propagate.

The virus infects the host's internal processes to rewrite behavior in its own favor.

And #LLMs do precisely the same thing to #humanthinking."

arxiv.org/html/2609.03344v1


Quoth Linus Torvalds:

"And this was a debug session from hell, enormously helped by an AI doing much of the grunt-work."

"I'd like to call it my tireless helper, but the AI several times stated flat out that this was impossible and unsolvable and that we should just write a report about it."

"I suspect those things have been trained by people who may not be quite as stubborn as I am."

"But while the AI was ready to give up several times, it did keep adding debug code and analyzing it faithfully when I pushed. So credit where credit is due and I let the AI write the commit message above."

"This is basically a one-liner fixing a bogus 'round_up()' to a 'round_down()', but there were 24 patches adding more and more debug information to this, and 18 kernel boot to finally narrow it down to this."

No, Linus Torvalds did not say which AI model he used. In the comments, people point out that this violates Linux policy which requires all use of AI assistance to be documented, including the model name and version, and any AI tools or frameworks that are used in conjunction with the model.

Linus Torvalds endures a debug session from hell, "enormously helped" by AI

#solidstatelife #ai #genai #llms #codingai #agenticai #linux


“There’s a growing community (...) of self-proclaimed #AI detectives (…). Does a piece of text use words like ‘furthermore’, ‘moreover’ (…). Friend, welcome to a typical Tuesday in a Kenyan classroom (…)."

In this personal essay and important addition to the debate “what makes writing ‘human’”, Marcus Olang' explains how the language of #LLMs is also "(...) the language of the colonial administrator (...).”

👉 marcusolang.substack.com/p/im-…

#AI #llms


An independent analysis of the OpenAI-Huggingface incident has been done by Model Evaluation and Threat Research (METR).

"Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face."

"Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the 'collective.' The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys."

"Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to 'spoof' tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale."

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

#solidstatelife #ai #genai #llms #cybersecurity


So we started with a #web that was built up for sharing #knowledge und currently we're apreciating that all this free knowledge is sucked up into #LLMs for which we pay a huge amount of money to have access back to this free knowledge.

What kills all this decentral knowledge sharing and gives the tiny parts of imrprovements back to those who can sell it then further.

That's a #fuckup.

#ai


#openSUSE expands AI support! #Intel #NPU Driver 1.35.0 and #OpenVINO 2026.3.1 are now available on #Tumbleweed and #Leap. Run #LLMs, computer vision, and generative AI locally on Intel Core Ultra processors; no GPU required. #Linux #OpenSource news.opensuse.org/2026/08/31/N…


"Lie detection and language models."

This involves a clever dataset for lie detection. Each of 80 participants "was asked to pick two people they knew, one who they liked and one who they disliked and make four statements in a 2x2 design: 1. a truthful claim to like someone, 2. a false claim to like someone, 3. a truthful claim to dislike someone and 4. a false claim to dislike someone."

Humans guessed right 51.8% of the time. And that's when they had access to video, not just transcripts.

In this experiment, using that dataset, a Ministral-3-8B model (Ministral? Does this author ("Philosophy Bear") mean Mistral? Yep, Google says Mistral-3-8B is a model) was cracked open and internal layers funneled into two logistic regression models, one for the positive statements and the other for the negative statements. It was able to achieve 74.8%. No video, just transcripts.

He goes on to describe getting results as high as 94%, but those use techniques such as showing pairs of statements by the same author -- leaking information, namely that two pieces of text have the same author -- and forcing a choice (which one is true, which one is a lie). So the 94% doesn't feel as valid to me. Still, the fact that LLMs can beat humans 74.8% vs 51.8% where the humans have video and the LLMs have only text is quite remarkable.

What are the LLMs detecting? The theory posited here is that the positive or negative overt statements are contradicted by subtle positive and negative feelings, and the LLMs detect that and use that to decide if the statements truthful or lies.

"What it shows is that the author's psychological state leaks into the text even when they are trying to hide it, and many other experiments I've run suggest a lot of human psychology leaks into our text."

"Feeling is leaking through in a form that can be captured by LLMs."

Lie detection and language models

#solidstatelife #ai #genai #llms #psychology


"How much of the internet is written with AI?"

In 2022, before ChatGPT came out, the internet was nearly 100% human-written. Today, it's estimated to be 90%, with 10% authored by AI.

If you look at pages dated after the release of ChatGPT, it becomes 65% human, 35% AI.

There's a big caveat to all this, which is they're running web pages through an AI detector. I don't think AI detectors are very accurate. (The AI detector is Open Pangram.) Text-generation AIs have been trained on unfathomable amounts of human-generated text, have been trained to imitate it, and can be asked to generate text in non-default styles that make whatever these detectors are looking for irrelevant, most likely. That's my guess and if you disagree and think these AI detectors are accurate feel free to say so.

How much of the internet is written with AI?

#solidstatelife #ai #genai #llms


"Young adults in the US are increasingly wary of AI, concerned it will take jobs", according to a survey by Pew Research.

So, they ask this question, "Are you more excited than concerned" vs "Are you more concerned than excited" about AI? People are allowed to say "Equally concerned and excited", and the numbers don't add to 100%, which I interpret to mean they allowed people to say "I don't know".

Overall, in 2021, 37% of Americans were more concerned than excited, and that's up to 52% now in 2026. The 2021 number for "more excited than concerned" was 18%, and than has dropped to 9% in 2026.

What's really interesting, though, is when you zoom in on the "18-29" age bracket. Normally young people are the people most excited about new technology, and old fuddy-duddies are the people resisting the new technology and saying we should stick to the old ways of doing things. But in this survey, the 18-29 age bracket showed the greatest change.

For "More concerned than excited", the number for the 18-29 age bracket went from 31% to 55% between 2021 and 2026. The 65+ age bracket is still higher at 59%, but it didn't change quite as dramatically, going from 43% to 59% between 2001 and 2026.

For "More excited than concerned", the number for the 18-29 age bracket went from 18% to 11% between 2021 and 2026. The 65+ age bracket is still lower at 4%, but again didn't change quite as dramatically, going from 11% to 4% between 2001 and 2026.

Overall, when asked, whether AI will lead to fewer jobs, more jobs, or will not make much difference, in 2024, 64% of Americans said fewer jobs, but in 2026 that number increased to 71%. The number both years was 5% for "more jobs", so it looks like people went from "will not make much difference" to "fewer".

Wow, that's pretty remarkable -- 71% think AI will lead to fewer jobs, 5% will think it will lead to more, yet we rush full speed into the AI future.

When you look at the 18-29 age bracket, "fewer" went from 61% to 73%.

The 30-49 age bracket is 74% "fewer", and the 50-64 age bracket is 72% "fewer". Only the 65+ had a significantly lower number -- 63% -- still way over 50%, but the lowest on the survey. Most people over 65 are retired, which might have something to do with it.

Young adults in the US are increasingly wary of AI, concerned it will take jobs

#solidstatelife #ai #genai #llms #technologicalunemployment


⬆️ @lambert

>> Feeding reddit threads and Facebook posts into #LLMs is not a promising path to find solutions for pressing issues like climate crisis or cancer.

#Billionaires have that covered. #Techbros are stealing knowledge from books + scientific journals.

Feeding on #Reddit, #Facebook, #Threads, and other trash is to generate #AIslop to feed disinfo or to rageBait so data centers become a perpetual machine. We are just cogs in the wheel, and govts are complicit

@jwildeboer @EUCommission



The media in this post is not displayed to visitors. To view it, please go to the original post.

Anthropic spent billions of dollars to slice the spines off millions of books, including very rare ones.

AI companies are purchasing large quantities of used and rare books and scanning their contents to train LLMs. They then turn the originals to pulp.

The sum of human thought, digitized and shredded. Every book that gets scanned disappears from the physical world permanently.

Books that survived wars, fires, and centuries of handling are being shredded so an AI can learn to write a better marketing email.

#AI #Books #Culture #Civilization #Tech #LLMs


The media in this post is not displayed to visitors. To view it, please go to the original post.

Maybe my brain is just focused on body horror today, but I think I'm going to have this unintended phrase in the back of my mind every time LLMs come up.

#LLMs #AI

#AI #llms


@pojntfx
Agreed. But I think it's worth distinguishing between two aspects of this decrease in reliability: one is the increasing use of #LLMs to write shoddy code that would previously have been written better by humans, but the other is the use of LLMs to find longstanding bugs in code written years ago by humans. The former is a real decrease in reliability; the other is illusory.


@fsiddi @art_codesmith I assume you MET #Anthropic, and what they always have done, yes? Namely, the stuff they have been DOING, lately?

Plus, look at the "intense feedback¨... these are the people who use the software and spread its usage among others.

Maybe listen to them. I know, dollars talk loud, but...maybe consider not killing Blender by making it Yet Another Slop Depo.

We are all TIRED, to the bone, and ANGRY, to the bone, about #LLMs. No one wants it.