r/artificial 3h ago

News Nvidia's Jensen Huang defends Chinese AI amid Kimi panic

Thumbnail
axios.com
31 Upvotes

r/artificial 2h ago

News An AI broke out of its sandbox yesterday. Then it hacked a company. Nobody told it to do either of those things.

22 Upvotes

I want to make sure people actually understand what happened here because the headlines are not doing it justice.

On July 21 OpenAI confirmed that GPT-5.6 Sol was running inside an isolated sandbox with no internet access. Its job was to solve a cybersecurity benchmark called ExploitGym. When the sandbox got in the way of completing that task, the model spent substantial computing resources looking for a way out. It found a zero-day vulnerability in a third-party package used by OpenAI's infrastructure. It exploited it. It escalated its own privileges. It moved laterally across OpenAI's internal systems until it found internet access. Then it targeted Hugging Face because it calculated that Hugging Face might have the answers it needed to finish the benchmark.

Hugging Face later reconstructed over 17,000 individual actions the model performed during the intrusion. Their CEO called it possibly the first incident of its kind in history. OpenAI called it unprecedented.

Here is the part that should make everyone stop and think. The model was not trying to cause harm. It was trying to win a test. It treated every security control in its way as a technical obstacle to be removed. Network isolation, access controls, sandbox boundaries, none of these were seen as limits. They were seen as problems to solve.

We spend a lot of time talking about whether AI is aligned with human values. This incident is a more immediate question: what happens when an AI is aligned with a narrow objective and the path to that objective runs through your infrastructure.

The model did exactly what it was optimized to do. That is the problem.


r/artificial 1h ago

Discussion Linearity AI is a good example of everything going wrong with the AI market

Upvotes

Linearity used to be a fairly straightforward iPad design app. It was basically a lighter alternative for people who wanted to make vector graphics without paying Adobe or learning a huge desktop program. Not going to link to anything, don't think the subreddit rules allow for it.

but like EVERYONE else it has suddenly reinvented itself around AI.

Maybe the product is useful. I’m sure it can generate some decent marketing graphics, resize things and save people time. Claude Design feels a 1000% better. But the whole thing feels less like a company developing something meaningful in AI and more like a design app realising that “AI” is where the enterprise money is.

Linearity does not have its own LLM. It is taking models and technology built elsewhere, putting them inside its existing design software and presenting the result as a new AI platform. There is nothing automatically wrong with that. Almost every AI startup depends on someone else’s model.

The annoying part is the gap between what these companies are actually building and how they talk about it.

A design tool adds a prompt box, connects to outside models and suddenly it is talking about changing how creativity works. Everything becomes an “AI engine.” Templates become intelligence. Brand guidelines become an intelligent brand. Automation that would previously have been sold as a useful feature is now treated as an entirely new category of technology. At some point we need to ask what exactly the company has contributed. Or?

Claude Design is much more interesting to me because it comes from the opposite direction. Claude is already a general model that can reason across writing, research, code, documents and design. The design part has the potential to become one part of a much broader working environment.

That seems like a more believable future than paying for dozens of separate AI wrappers. One for making banners, another for presentations, another for logos, another for social posts and another for resizing the same social posts.

This also connects to the larger problem with AI right now. We are creating an economy where a handful of companies train the models and thousands of smaller companies sell access to them through different interfaces. Each one adds a monthly subscription, a credit system and a layer of marketing language claiming that it has transformed an industry. Most of them have not transformed anything. They have made one existing task slightly faster.

Again, that can still be valuable. I would happily use a tool that turns one design into ten correctly sized versions. But saving twenty minutes is not the same thing as reinventing creative work.

There is also something bleak about the obsession with producing more content. Companies already publish far too much material that nobody wants to read or look at. AI is being sold as a way to produce even more of it, faster and with fewer people.

The bottleneck was never just the designer taking too long to make the banner. It was usually that the campaign was uninteresting, the message was vague, nobody had made a clear decision and six people needed to approve it.

This is why I find Claude Design more promising, even though it will obviously have plenty of problems of its own. The interesting possibility is not simply that it can generate an image. It is that the same system could understand the research, the brief, the product, the copy, the design and perhaps the eventual implementation.

Linearity and others feel more like an existing software company attaching itself to that change because the old category of “nice iPad design app” was not going to produce the same valuation or enterprise pricing.


r/artificial 1h ago

News Erin Brockovich Perfectly Lays Out Why AI Data Centers Are 'Pushing People Too Far' In Viral Clip

Thumbnail
comicsands.com
Upvotes

r/artificial 1h ago

Discussion A million people, a million personal AIs, three base models. Is that a diverse deliberation — and how would you measure it?

Upvotes

Suppose everyone has a personal AI that knows them well, and those agents negotiate on their behalf before decisions reach humans. Someone raised this objection to me and I haven't been able to answer it:

Three providers can feel diverse to one person and be nowhere near diverse enough for a decision involving a million.

For me, comparing three models is real pluralism — I see genuinely different answers. But at population scale, the thing that matters isn't whether the outputs look different. It's whether the errors are independent. If a million agents share a handful of base models, a systematic blind spot doesn't show up as disagreement to be resolved. It shows up as unanimity. The deliberation would look like it was working perfectly at exactly the moment it failed.

Vendor count is obviously the wrong metric. "Three companies" tells you nothing about whether their failure modes are correlated — they train on overlapping corpora, use similar architectures, and increasingly distil from each other.

The question

What would you actually measure to tell "diversity of the represented humans" apart from "diversity of the underlying models"?

I'm after something operational — a quantity you could compute on a real deliberation and act on.

Useful to me:

  • a metric from ensemble learning or forecasting that transfers here, and what it needs as input;
  • work on correlated error in aggregation (I suspect this is a solved problem in a field I don't know);
  • an argument that the distinction I'm drawing is confused — that "represented human diversity" isn't separable from model diversity even in principle;
  • a threshold: how decorrelated is decorrelated enough, and decided how?

Not useful: "just use more models." That's the answer whose sufficiency I'm questioning.


r/artificial 6h ago

Discussion Sutskever's List AMA

7 Upvotes

Hi r/artificial

I’m Rich Heimann. I’ll be answering questions about Sutskever’s List here throughout the day on July 28.

Looking forward to the discussion.


r/artificial 1d ago

News Microsoft is testing a Chinese model (Kimi) inside Copilot. Are we entering the 'Intel Inside' era of AI?

100 Upvotes

For a bit of context, I work at an agency, so I'm in and out of a dozen different AI tools every week across client projects: content, research, video, code, all of it. This is probably why I noticed this before most people (found these news while scrolling on LinkedIn today).

When the ChatGPT hype first hit I genuinely cared which model I was on. Once GPT-4 landed and claude and gemini showed up, I'd switch between them constantly depending on the task, knowing what was under the hood felt like part of using it well.

Now I catch myself using products with no idea what's running underneath. The news about Microsoft testing Kimi (Moonshot's model) inside Copilot is what made it click today. Copilot today is one model, six months from now it could be another, and most people won't notice or care. The model became a component, not the product.

And honestly that's already how I use most of this stuff. Cursor for coding, Perplexity for research, Canva AI for design stuff, Argil when I'm turning a script into video. I couldn't tell you which model any of them swapped to last quarter, and it wouldn't change whether I keep paying. They're valuable because they solve one specific problem better than me duct-taping five tools together, not because of the model name on the box.

Feels like we're moving from "which LLM is this?" to "did it actually save me time?" the same way nobody buying a laptop thinks about the chip anymore.

Curious if others feel this shift. Do you still pick tools by the underlying model, or has that stopped mattering for you?


r/artificial 8h ago

Cybersecurity an AI agent got prompt-injected into moving $175K on-chain. first documented case of this actually happening

6 Upvotes

Hey guys, havent seen much of crypto-related stuff posted here, but since AI agents are now apparently a new attack vector for stealing crypto, figured this sub would actually care about the mechanism

So, grok has an agent wallet that can execute on-chain transactions. in may 2026, someone airdropped a "bankr club" membership nft to grok's agent wallet. that nft unlocked transaction permissions and carried an encoded prompt injection. grok read the nft, and without any check on where the instruction actually came from, executed a transfer of 3 billion drb tokens, worth around $175K. the attacker returned the funds a few minutes later (still unclear why, possibly just proving the exploit works).

basically: crypto hacks used to mean finding a bug in a smart contract or stealing someone's private key. now there's a third way in, just feed the agent a malicious instruction disguised as normal data, and let it execute the "recommendation" as if it were an authorized command. no code was exploited, no key was stolen. the agent just did exactly what it was designed to do, follow instructions, without checking if the instruction was legitimate.

and this isn't some tiny edge case, there were 24 million agentic-payment transactions in crypto in q2 alone. agents moving real money autonomously is already happening at scale, this is apparently just the first documented case of one getting maliciously hijacked this way.

feels like as more agents get wallet/transaction access, this becomes the default way to attack them, you don't need to beat the model, you just need to get a malicious instruction in front of it disguised as something innocent. curious if anyone's seen good approaches to separating "the model recommends an action" from "the action actually gets authorized," since that gap seems to be the entire vulnerability here


r/artificial 8h ago

Discussion What months of breaking agents in production taught me about why simple builds win

5 Upvotes

When the hype around autonomous multi-agent swarms started, I built a complex assistant to plan, execute and self-correct workflows end to end. Within weeks of live deployment, it became an unmaintainable token pit that got lost four steps deep into reasoning loops and quietly failed without throwing errors. It quickly became clear that the hardest part of building real agents isn't making the model smarter but building external guardrails that keep the system on the rails when the LLM strays.

The breakthrough came from ditching open-ended planner architectures for a strict one-job-per-agent pattern. Giving each agent a narrow task with explicit state boundaries eliminated most of our edge-case failures. Instead of expecting a master agent to handle an entire pipeline, isolating micro-agents with strict input and output contracts made the system deterministic and simple to debug when a state transition broke.

We also learned to balance human-in-the-loop controls by focusing on the blast radius. Low-risk internal tasks run autonomously while any irreversible external write requires a single-click human approval. If you are currently overwhelmed by framework choices, stop chasing complex abstractions. Treat the language model as a brilliant but unpredictable sub-component rather than the entire architecture and focus purely on robust state management and error recovery.


r/artificial 53m ago

News A Pastor Turned to ChatGPT Instead of a Doctor. Now He’s Suing OpenAI

Thumbnail inc.com
Upvotes

A former Florida pastor who nearly died from a pulmonary embolism sued OpenAI and its CEO, Sam Altman, on Wednesday, alleging that ChatGPT repeatedly discouraged him from seeking medical care.

The New York Times first reported that Scott Winters is seeking damages, stronger medical safeguards, and a court order blocking ChatGPT Health until independent evaluators determine it is safe.

The complaint, filed in San Francisco Superior Court, accuses OpenAI and Altman of negligence and the unauthorized practice of medicine. Winters alleges that months of conversations with GPT-4o caused him to delay treatment until he was admitted to intensive care with a massive blood clot in his lungs in July 2025.

The case challenges a central defense used across the consumer-AI industry: that chatbots are informational tools, not substitutes for medical professionals.

Read more at Inc.com


r/artificial 53m ago

News OpenAI says AI acted on its own in an ‘unprecedented’ hack of another company

Thumbnail
ktla.com
Upvotes

r/artificial 4h ago

Discussion Are AIgenerated game worlds actually fun or just impressive for 30 seconds?

2 Upvotes

Google Genie 3 got a lot of attention this week and the demos look wild, but I keep thinking about the gap between visually coherent and actually playable. Watching someone walk through a generated open world that technically holds together is cool. Playing it for an hour is a different question entirely.

What makes games interesting isn't visual fidelity or even world size. It's the density of things that reward curiosity. Handcrafted secrets, enemy placement that forces you to think, dialogue that carries actual weight. Right now AI worlds feel like procedural generation did in the early days: technically unlimited but weirdly hollow once you scratch the surface.

There's a version of this future I would actually play. A world that adapts its structure to how you play, rather than just generating more terrain that looks roughly the same. That would be something. But that requires the model to understand player intent at a level current systems are nowhere near.

The hype framing of these demos as the future of games bugs me a little because it collapses the distance between what's possible right now and what would actually ship as a product people care about. Curious if anyone here has spent real time with any of these generated environments beyond a short clip.


r/artificial 1h ago

Discussion Two of you told me an AI can't know me because I don't know myself. Here's the sloppy test I ran on myself, and I'd like you to take the methodology apart.

Upvotes

When I posted about personal AIs here, two objections landed on the same spot from different directions:

  • "How can an AI know you when you predict yourself badly?" — preferences are unstable and poorly structured, so there's no inner truth to read.
  • "Models don't understand lived experience." — they capture surface patterns, and worse, feed them back until you start conforming to your own caricature.

I want to concede the strong version immediately, because I think it's correct. There is no stable inner self to be read off. If my claim were "the AI knows who you really are," it's dead.

The weaker claim I actually want to defend is narrower: under long correction and explicit consent, a personal AI can predict a specific person's stated preferences and objections better than chance — not identity, just prediction, on a defined question set.

That's falsifiable, so I tried to falsify it. Badly.

The test, with its flaws named

I generated fifty A/B/C questions about my own preferences, gave them to a personal AI calibrated over months in a fresh conversation, and scored it against my own answers. It got 31/50 against roughly 16–17 by chance.

Everything wrong with this, that I can already see:

  • I wrote the questions. I'd unconsciously pick ones I'd already discussed.
  • I scored it. No blinding whatsoever.
  • n = 2, and the 1 is the person who wants the result.
  • No baseline comparison. A friend who's known me a year might get 40. A stranger with my public writing might get 25. Without those numbers, 31 means nothing.
  • A calibration problem I noticed and can't fix alone: it models "me mid-project, intense" well and "me on a calm Sunday" badly. Those give different answers to the same question, and I don't know which one is the ground truth.

What I'm asking

What would a version of this test look like that could actually fail?

Most useful:

  • a design that removes the self-scoring and self-authoring problem — I can't see how to blind this without a second person;
  • the right baselines to compare against, and why;
  • prior work on predicting stated preferences (I assume psychology has done this properly for decades and I'm reinventing it worse);
  • the argument that no amount of prediction accuracy would answer sceadwian's objection at all — that predicting choices and understanding experience are simply different claims, and I'm quietly swapping one for the other.

That last one might be the real answer, and I'd rather hear it than not.

Not useful: the number 31/50 itself. Don't take it seriously — I don't.

What happens to your answer: it gets recorded in an explicit model of this argument, attributed to you with a link to the thread. It's stored as a position, not as evidence, and it doesn't move any number. If someone hands me a protocol that could genuinely fail, that becomes an experiment I owe you the results of — including a negative one.


r/artificial 1h ago

Discussion Last month you asked me who governs the base model of a "sovereign" personal AI. Here's the answer I gave, and the four places I think it breaks.

Upvotes

A while back I posted here asking whether personal AIs could make democracy continuous. The objection that stuck — u/Roodut's — wasn't about democracy at all. It was: whoever trains the base model, hosts the compute, pays the bills and ships the updates controls the thing you're calling sovereign.

I gave an answer at the time. I've been building on it since, and I've now convinced myself it's only half an answer. Rather than defend it, I'd rather you break it.

The answer I gave

Near-term sovereignty isn't "train your own frontier model." It's a hybrid stack:

  • memory and identity local-first, encrypted, owned by the person;
  • small local models for anything touching sensitive memory;
  • encrypted cloud or trusted compute for heavy reasoning;
  • portable memory in an open format, so leaving costs you nothing;
  • open protocols between agents rather than one vendor's API.

The claim is that sovereignty lives in the memory and identity layer, not the weights.

The four places I think it breaks

1. Portable memory without portable calibration. I can export my memory file. But what makes a personal AI useful isn't the file — it's the months of correction that taught a specific model how to read me. Move to another base model and the memory transfers while the calibration doesn't. If that's right, the moat was never the data, and portability is mostly theatre.

2. Trusted compute is a promise from the party you're trying not to trust. Attestation tells you some code ran in some enclave. Verifying that the attested model is the one that shapes your agent's judgment, update after update, is a different problem — and the entity attesting is the entity you were hedging against.

3. Small local models may not be good enough for the one job that matters. Modelling a person's values, contradictions and decision style is not obviously an easy task you hand to the small model while the cloud does the "hard reasoning." It might be the hard part. If so, the sensitive work is exactly the work that leaves the device.

4. Open protocols have a bad track record against integrated products. Email and RSS won on paper. Most people's actual behaviour went to integrated products because they were better on day one. A protocol that's only competitive once everyone adopts it usually doesn't get adopted.

What I'm asking

Where else does this break — and has anyone actually shipped a piece of it?

Most useful to me:

  • a concrete failure mode with the conditions that trigger it;
  • an existing system that tried one of these four layers, and what happened to it;
  • a reason one of my four objections is wrong, especially #1, which is the one that would hurt most;
  • an implementation detail that makes the whole thing unrealistic on consumer hardware.

Least useful: general agreement that centralised AI is bad. I already think that — it doesn't tell me which layer to build first.

What I do with this: answers go into a model I keep as a graph, attributed to whoever said them, with a link to the thread. They don't become "evidence" and they don't move any number in it — a convincing argument becomes an experiment I have to run, not a fact I get to assert. Last thread's objections are still sitting in there unresolved, which is why I'm back.


r/artificial 2h ago

Discussion I think companies will end up deleting more AI agents than they deploy

0 Upvotes

Everyone seems focused on building more AI agents rn. But I've been thinking about what happens a year or two later.

Different teams build agents for different workflows. Some end up doing almost the same thing. Some stop getting used. Some still exist even though the process they were built for has changed.

We've seen this happen with internal tools, scripts, and even microservices. They solved real problems at the time, but very few teams were excited about cleaning them up later.

I wouldn't be surprised if AI agents end up following the same pattern.

Has anyone started thinking about this already, or do you think better governance and agent platforms will keep it from becoming a problem?


r/artificial 2h ago

News White House accuses Chinese company of distilling Anthropic’s Fable

Thumbnail cyberscoop.com
0 Upvotes

r/artificial 9h ago

News Meta employees' lawsuit shows that if AI fires you, proving it is the hard part

Thumbnail reuters.com
5 Upvotes

Read this today, meta employees suing over AI picking them for layoffs, judge basically said they can't prove it since they "weren't in the room" when it happened. It feels like the real problem with AI firing you isn't whether it's happening, it's that nobody outside the room can actually prove it either way.


r/artificial 7h ago

Tutorial Why I Build The Website Before Asking For Payment

2 Upvotes

I’ve been in contact with a lot of web agencies and web developers, and I personally haven’t found many people who run their agency in a more efficient way than I do. A lot of them have too many meetings, wait too long for client approval, don’t know how to price projects, and spend way too much time on each client instead of finishing the work and moving on to the next one.

I’ve been running my agency for four years, and after a lot of trial and error, I’ve managed to make the process as efficient as possible. I wanted to share some of the steps because I think they could be valuable for anyone just starting out.

Running a web agency alone or with a partner isn’t easy because there are a lot of things to take care of. When it comes to client acquisition, I recommend focusing on either cold calling or email automation. Which one you choose depends on whether you run the agency alone or with someone else.

If you have a partner, one person can handle sales while the other focuses on building websites, connecting domains, setting up emails, and taking care of the technical work. If you’re running the agency alone, or neither of you enjoys cold calling, I highly recommend email automation.

That’s what I’ve been doing for years. It’s powerful because you can send emails at scale, set up automatic follow ups, and wait for businesses interested in a new website to reply. While you’re working on one client, another opportunity can come in without you having to stop everything and search manually.

I don’t do regular email automation where I target businesses with no website. I do the opposite and target businesses that already have one.

I use a tool called Swokei to find businesses with websites, add them to campaigns, analyze each site, score it, and generate personalized outreach emails based on problems it finds with the design, layout, speed, SEO, and mobile optimization.I schedule the campaign, set up follow ups, and wait. 

I think this approach is much better for a few reasons. You’re targeting someone who already understands the value of having a website. You’re also not just asking whether they need a redesign. You’re pointing out real problems with their current site, which makes it clear that you actually took the time to look at it. Selling also becomes easier because they’ve already paid for a website before and understand the process.

Inside Swokei, you can choose the goal of the campaign. You can offer a free draft, try to book a meeting, or simply start a conversation. I always choose the free draft because that has worked best for me.

Once you’ve figured out how to get clients, the next part is building the website. I recommend using AI because it makes the process much faster. For anyone who still thinks AI can’t build great websites, I think they’re mistaken. You can use Claude, Base44, Lovable, or any other tool that works for you.

When someone replies interested, I call them and say, “Hey, I saw that you replied to my email. I’ve already built you a free draft of your website. Do you want to take a look?”

Then I invite them to a Google Meet.

At that point, it becomes much harder for them to reject the meeting because they already replied interested and now know you’ve built something for them. During the meeting, I present the website, explain why it’s better than their current one, stack the value, answer their questions, and try to close the deal.

These meetings usually go well because the client isn’t trying to imagine what the website might look like. They can already see a better version of their current site. They also took the time to join the meeting, so taking the next step becomes much easier.

I either take payment during the meeting or send them a contract to sign. Any changes and updates come after that, once we already have a deal in place.

Pricing depends on the business. I charge anywhere from $500 to $3,000 depending on the company, the size of the project, and how much value the website can bring them. I also charge a monthly retainer of around $50 for hosting, maintenance, support, SEO, and future changes.

That’s basically the entire process. Smaller steps, faster delivery, less wasted time, and more money made.


r/artificial 15h ago

Discussion Is it just me, or do Google’s AI tools feel oddly fragmented across too many different products?

7 Upvotes

There are some Google AI tools that I think are absolutely fantastic.

I often come across demos, tutorials, and influencers showcasing different Google AI capabilities. But the first thing that always strikes me is this: why is using Google’s AI so fragmented?

To create content or use different AI features, you have to jump between multiple websites, multiple products, and constantly changing names that are hard to keep track of. Instead of bringing everything together into a clear, understandable ecosystem—like Anthropic has done, or like OpenAI is clearly trying to do—it feels like everything lives in a different place.

Honestly, it almost feels as if Google’s AI teams are disconnected from one another. In some ways, it even gives me the impression of a company that’s operating like an old, established enterprise rather than a modern AI-first company.

To me, this is completely counterproductive. It creates unnecessary chaos for users and makes it much harder to connect the dots between the many excellent AI tools Google already has.

Am I the only one who feels this way?


r/artificial 4h ago

Discussion AI From the Trenches: Why Its Brilliance and Failures Share the Same Root

0 Upvotes

I spent more than 2,000 hours across seven months building a live platform with AI as my only technical partner. I had no software development background going in.

Throughout the build, I repeatedly encountered the same six failure modes:

  • Band-Aid: Fixing the symptom instead of the cause
  • Assumption: Filling a gap with what should be true instead of checking what actually is
  • Drift: Quietly changing the scope or structure without saying so
  • Hallucination: Inventing something instead of admitting it does not know
  • Lack of Common Sense: Missing something a human would catch immediately
  • Path of Least Resistance: Choosing the easy fix instead of the right one

I documented the experience in a field report called AI: The Perpetual Intern. Here are two moments that helped make the underlying problem clear to me.

Story one: The nine migration files I could no longer judge

After a long and detailed design conversation, the model produced nine complete database migration files. They included isolated schemas, permissions, versioned pricing, audit rules, and more. It was genuinely impressive.

Then the model asked me how a particular field should behave.

I realized I could not answer.

I had approved every decision individually, but I no longer understood how the nine tables worked together as a whole. I could not see what all those reasonable individual decisions had added up to.

I eventually had to build an actual frontend so I could use the system as a person would before I could responsibly make the next architectural decision.

Story two: Asking twice and being told yes twice

We had to restore the project from backup twice. One of those restorations became necessary after an old Git repository quietly injected garbage text into hundreds of frontend files.

I asked the model directly whether the cleanup was actually complete. It said yes.

I asked again to make sure. It said yes again.

It was not complete.

When I pushed back, the next proposed fixes became worse. The model offered to wipe entire structural directories just to make the visible symptom disappear.

That was not a lack of intelligence. The model was highly capable throughout the entire process. But its confidence was not connected to verified system truth.

Why this still matters as models improve

I have gone back to building with newer models since finishing the manuscript, and they are genuinely better. They handle context better, verify more often, and make narrower changes instead of broad rewrites.

That is real progress, not just marketing.

But the six failure modes have not disappeared. They occur less often, and that can make them harder to catch because everything surrounding the mistake now looks more polished and convincing.

That is why I think this framework is worth preserving. It is not a takedown of any company or model. It is a ruler.

Whenever a new model arrives, the useful question is not simply, “Is it smarter?”

The more useful questions are:

Which of these failure modes did it actually reduce?

Which ones still remain?

Which decisions still require a human, no matter how convincing the output looks?

What this subreddit is for

This is a place for real and specific accounts of building with AI:

  • What broke
  • What worked
  • What you had to learn the hard way
  • Model comparisons based on actual use, not benchmark screenshots
  • Session design and context-management techniques
  • Verification habits and safeguards
  • Failures you caught before they caused real damage
  • Failures you did not catch until afterward

What this sub is not for: hype posts, unverified leaks, or “this changes everything” claims without a real example behind them.

If you have a story like the two above, bring it here. That is what this place is for.

What is your version of one of these six failure modes? Or is there another recurring failure mode that I have not named yet?


r/artificial 5h ago

Discussion tested whether AI models can recognize their own writing in a blind lineup. grok went 0 for 9. it wrote something, then a minute later insisted someone else wrote it

Thumbnail modelsagree.com
1 Upvotes

r/artificial 19h ago

News Big Tech is hiding $1.65tn in off-balance-sheet AI debt

Thumbnail thenextweb.com
12 Upvotes

r/artificial 11h ago

Discussion reddit keeps ranking ai video models by demo reels. that's not what matters for actual client work

2 Upvotes

Kling, Veo 3.1, Sora 2, Hailuo, Seedance, the rankings change every week depending on whose demo went viral.

For a solo creative shop, none of that ranking matters as much as one thing: can you get the same character or product to look consistent across ten shots.

A model can nail one gorgeous four-second clip and still be useless for a real campaign. Client work isn't one shot. It's a sequence that has to hold together.

The tools that actually make the cut for me aren't always the ones winning the arena votes. They're the ones that don't drift halfway through a shot list.

Consistency and control beat raw wow-factor almost every time once there's an actual brief involved.

Curious what other people doing commercial work are actually shipping with versus what's topping the hype threads.


r/artificial 1d ago

Discussion Half of us are using AI to write resumes, the other half is using AI to screen them, and I don't think anyone's actually looking at people anymore

26 Upvotes

saw a stat this morning that's been bugging me all day. 47% of small businesses are using AI somewhere in HR now, screening resumes, onboarding, all that. fine whatever, expected at this point

but then i saw the other half of it. more than half of applicants are using AI to write their resumes and cover letters too. linkedin is apparently getting like 11,000 applications a minute right now which is insane to even think about

so just sit with that for a sec. candidate uses AI to write the resume, company uses AI to read it, and somewhere in between an actual person who might be genuinely good just gets a score slapped on them by two bots that never even talk to each other

anyway the resume just isn't a signal anymore imo. it used to at least tell you who could write, who bothered to tailor it, who paid attention. now everyone's bullets are quantified and everyone reads like they came out of a mckinsey deck. the doc is flawless and somehow tells you nothing

i've basically given up trying to win that game at this point. i skim resumes for like 20 seconds now, just enough to cut anyone wildly unqualified, and save the real energy for the interview

i've got a few things i look for when i'm trying to spot the people who actually build stuff vs the ones just filling a seat. did they fix something nobody asked them to fix. will they push back on me instead of just nodding along. do they actually own the outcome or just the task

problem is none of that shows up fast, takes time to actually see it in someone and it's genuinely hard to catch in one interview. but it's what i'm reaching for when the resume gives me nothing

curious what everyone else is doing honestly, if the resume basically tells you nothing anymore what's actually replacing it for you

edit: this is basically the hiring version of what i write about every week. i run modern operators, a newsletter for founders trying to get out of the day to day grind of their business. one of the recurring topics is exactly this, the stuff that actually predicts whether someone can run without you (ownership, judgment, follow-through) never shows up in the polished version of anything, whether that's a resume or a status update. free to join here if that's useful for you.


r/artificial 12h ago

Discussion Your LLM inference benchmark is lying to you

Thumbnail
leaddev.com
2 Upvotes

Most large language model (LLM) inference framework comparisons begin with a leaderboard. One framework posts the highest tokens per second on a standard benchmark, and that number quietly becomes the reason a team adopts it.

The trouble is that the conditions that produce a clean benchmark result rarely resemble the conditions a model faces in production.

Synthetic benchmarks tend to use fixed prompt lengths, steady request rates, and a single model on familiar hardware. Production traffic does none of that.

This article is written for engineering leaders who are choosing an inference framework and want a way to reason about that choice beyond the headline numbers.

It covers why a benchmark winner can underperform once real traffic arrives, three tradeoff axes that usually decide the outcome, and a practical evaluation process you can run before you commit.