r/slatestarcodex 21d ago

Monthly Discussion Thread

6 Upvotes

This thread is intended to fill a function similar to that of the Open Threads on SSC proper: a collection of discussion topics, links, and questions too small to merit their own threads. While it is intended for a wide range of conversation, please follow the community guidelines. In particular, avoid culture war–adjacent topics.


r/slatestarcodex 21h ago

Links For July 2026 (Part 1)

Thumbnail astralcodexten.com
18 Upvotes

r/slatestarcodex 8h ago

Rationality How Much Do You Value Online Anonymity?

19 Upvotes

Read this insightful blog piece about anonymity: https://sive.rs/anon and have been debating the merits of full vs partial vs zero online anonymity. On one hand it's nice having the confidence that my digital footprint can't be traced back to me. But that's probably too rosy of an ideal when data brokers and companies like Palantir exist. The history of Scott's blog is if anything a data point against the possibility of true online anonymity. One big advantages of zero anonymity is having the ability to liquidate all the social capital you've built up. Having a name and a face to a brand or blog is so powerful in selling that image. Scott's probably a perfect example of this where I presume he only got more and more high value connections post being doxxed. Juxtapose this with someone like Gwern who said himself on one of dwarkesh's podcasts that he only makes a couple thousand a month (or maybe less I forget). If he were to "reveal" himself I'd almost guarantee he'd gain in social standing and the rest.

Anyways, curious on the community's thoughts and/or any links to some old slatestar blogs that might've touched on this


r/slatestarcodex 11h ago

Are there any dating apps (or other venues) that are popular among rationalist-adjacent folks?

15 Upvotes

I would like to date more rationalist-adjacent people because some of the more positive experiences I've had in the dating scene have been with folks like us.

For example, I went on a date with this guy a couple months ago, and he seemed nice but I hadn't clocked him as rationalist-adjacent. That's okay because I'm an equal-opportunity dating partner. But when he unexpectedly mentioned "I Can Tolerate Anything Except the Outgroup" on our date, my jaw dropped and I was immediately smitten for him.

In a similar vein, earlier this year I had an overnight tryst with someone who talked very knowledgeably with me about existential risk and gradient descent and other such topics. He was strictly casual so, in retrospect, I suppose I had been playing the role of one of those AI-fluent Silicon Valley escorts (though I wasn't aware of the phenomenon at the time).

Anyway, is there a dating app that rationalist-adjacent folks are particularly drawn to? OkCupid is a far cry from how it used to be so I can't imagine there anymore. I thought the women-first approach to Bumble might appeal to some of us, but women don't actually (need to) make the first move on Bumble so forget that.

Anything I should put in my profile to signal this? I have a prompt about p(doom) but the majority of responses to it are either: (1) people who don't know what it means and need me to explain it to them, or (2) people who think that AI is a useless text predictor that was built solely to line the pockets of billionaires and destroy the environment and then proceed to lecture me on that 🥲

Alternatively, is there anywhere in real life that hosts a disproportionate number of rationalist-adjacent folks?

Sorry in advance and please delete if inappropriate. The monogamy thread from last week got me thinking.


r/slatestarcodex 20h ago

Are Therapists to Blame for the Rise of Adults Cutting Off Their Parents?

Thumbnail msn.com
55 Upvotes

r/slatestarcodex 5h ago

Should robots be considered as slaves?

Thumbnail optimallyirrational.com
2 Upvotes

r/slatestarcodex 21h ago

The USA is facing an interesting Game Theory problem around Daylight Saving Time

18 Upvotes

Once again the USA is considering removing Daylight Saving Time. I'd like to write a short blurb about the game theory, the current politics/arguments, and why everyone reasonable should push for ANY action. I'm not a great writer, so I would welcome someone experienced like Gwern, Zvi, etc to expound on the topic in a more catchy way. I am firmly in the camp of "We should do something to solve this stupid game".

1. The Game Theory.

Y'all love game theory. Prisoners dilemmas, trading jelly chips in glowfics, robot negotiation programs.

Right now, there is a simple prisoners dilemma-type thing in real life:

Coordinate for Keeping DST: some people win, some people lose.

Coordinate for keeping standard time: Some people lose, some people win.

Do not coordinate: Most people lose.

Before getting into the details, this is an interesting game theory because of the sheer number of people involved and the nature of public discourse. If 50,000,000 people stand to lose, but 250,000,000 stand to win, will the 50,000,000 people feel strongly enough to be loud about their dissatisfaction to stall progress / lead to not coordination?

2. History, The current politics / arguments

1942-1966 - DST year round.

1966 - the Uniform Time Act - DST End of April to End of October

1973 - Nixon tried 16 month trial of DST year round. It was stopped early.

1986 - Made DST Longer - Start of April to end of October

2005 - Made DST Longer - Mid March through early Novemeber.

2022 - Permanent DST "rushed" through congress, never went to House.

2026 - Permanent DST slower through House, currently called "DOA" in congress.

Key arguments/debates so far:

For taking any action:

Switching sucks. The Majority of scientists and studies, along with anecdotes (especially among parents of smaller children) is that shifting clocks sucks. The president has laid out the question as 50% of people want standard time, 50% of people want DST, but everyone hates changing clocks. The key studies suggest that periods after shifting clocks leads to increased traffic deaths, suicides, poor learning outcomes of children, and notable other issues. Not taking any action is a net negative for almost all parties as the above issues with changing clocks is more significant than other issues listed below, and most people agree about this. I will mention that when I see facebook comments, around 1/20 comments suggest keeping the shifting clocks. The principle of "switching sucks" was mentioning in the house hearing several times, by the president, and seems RELATIVELY universally agreed.

For Keeping Permanent Standard Time:

Safety: Morning Commutes having more sunlight could decrease auto accidents. Children may wait for the bus in the dark less - the theory is that children waiting for the bus would be safer with light from the sun, to have fewer child car injuries.

Solar Noon: People argue that the sun should be the highest point when the clock shows 12, out of principle.

Sleep and Sun: People enjoy waking up around sunrise, but more importantly, people don't want to go to bed when the sun is still up. If people go to bed at 10pm, and the sun hasn't set yet, people strongly dislike that.

Absurdity: Moving clocks is silly when we could just schedule work/school/etc an hour earlier for a similar effect as DST. "You can't cut one end of a scarf and add to the other end to make a longer scarf".

Insanity: It was tried in 1973 and didn't work.

For Making Daylight Savings Time Permanent:

Energy: Generally, DST saves energy. I believe the main source of this is fewer lights being turned on in the evening.

Recreation / Health: By taking a morning hour (Most people are working/in school) and putting it in the evening, people are more likely to Exercise, recreate outdoors, and leave their homes.

Economy: Similar to above, when it is light, people are more likely to spend money on food, recreation, in stores, etc.

Changing Times: When it was tried in 1973, we didn't have computers, internet, phones, we had way fewer lights, more farmers, etc. For instance, over 50% of school children are now dropped off for school. Schools start than in 1970s, so many students are going to school in the dark even in standard time. DST would make no further impact to those students who are going in the dark already or are being dropped off.

Anti-Absurdity: It is easier to change the clocks than change school start times, work start times, etc. We have known that children do poorly on early schedules for decades, and yet almost no school district in the USA has shifted to a later schedule. This is because school start times vaguely follow work start times.

How should we actually choose?

A few ideas:

  1. Go with the current proposal. It's already passed one body of congress. Let's just try it and see, this is likely the fastest option.

  2. Do the math. There are many sites that start to help with this (like https://observablehq.com/@awoodruff/daylight-saving-time-gripe-assistant-tool), but I'm not sure any try (even with approximations) to answer the FULL question. Many people by location and preference might prefer one over the other. How many people would benefit from each method? How does it translate into QALYs? How many students in the areas would actually be impacted? How many QALYs would that impact? Lay out methodology and findings and we can rally around that answer.

  3. Do the math, but different. Apparently, we have already extended DST twice. Were studies done in the weeks adjusted to see if any of the above arguments hold more/less true? Have other changes in state time zones seen a specific change to commute-related accidents, outdoor activity, energy use, etc? Maybe we can use that data to further inform #2 and come together with findings we can rally around.

  4. Continue to lose the game theory.

I am a proponent of #1 - and if it fails, I will be a proponent of the next time this idea gets floated around. Thanks for reading! I will be happy to incorporate any further information anyone posts in the comments.


r/slatestarcodex 20h ago

I made a daily 5 question calibration tool in the style of the Scout Mindset calibration exercise

5 Upvotes

I’ve recently been thinking a lot about personal calibration.  Some of you may remember this from Julia Galef’s book “The Scout Mindset”.  If not, the basic premise is if your confidence in an answer matches your accuracy actually answering.  Essentially, Are you 90% right when you are 90% confident.  I found the questions from her calibration exercise interesting, especially as a fan of trivia. Many of her questions seem like general trivia repurposed as calibration.  There was even a now expired tool for answering Galef’s questions and scoring passed around here in the past:  https://www.reddit.com/r/slatestarcodex/comments/mtf6gy/scout_mindset_calibration_practice_this_is_a_tool/

It can definitely be argued how useful honing your calibration on general trivia can be, but I thought it would be fun to turn this concept into a sort of daily ritual.  Answer 5 questions each day and build a data set of confidence vs accuracy.  At best I might know myself a little better, at worst at least I might learn some interesting facts along the way.  It’s all disguised as a game that might actually convince a normal person to use it.

So I set about building the thing which led to some decisions which I’m interested to know other’s thoughts on.  

Questions are true/false like the original.  

Scoring is Brier based (p on the true outcome -1)^2.  

I display this (1 - Brier Score) x 20 so a 5 question round is out of 100, a nice round number. 50% confidence yields 15 points a question so 75 per round of pure guessing.  

I chose to do 6 confidence stops 50/60/70/80/90/99.  99% and wrong scores 0, correct scores 20.  These scores are all rounded to whole numbers for display but not in the data set.  I keep the raw figures so values don’t drift over time.

It’s a marathon, not a sprint.  5 questions a day takes a while to bank useful data.  I don’t even show the calibration chart until you have 40 questions in the can. This is the classic accuracy vs confidence bucket chart y=x diagonal meaning perfect calibration.  I also break this down by category.  So I could see I’m overconfident in science but under confident in history.

The questions come from an LLM.  I thought this would be the easy part “Hey claude give me 100 true/false trivia questions”.  Not at all.  I ended up with an entire pipeline which generates the questions (Gemini), an automated critic (Claude) evaluates each statement and auto rejects ambiguous ones that just don’t have a good calibration signal (widely known, unsurprising, etc), and flags / drops items whose model provided sources don’t support or only partially support the claim.  At the end of all that the questions all land in a queue for my personal review.  The yield is roughly 30-50% good usable questions.  

The app is called Hedge: Calibrated Trivia.  It’s free, no ads, no accounts or tracking.  No in app purchases, nothing like that.  iOS only (sorry). It’s just my solo passion project.  I’m a software engineer and I write about the project at https://stiles.one/hedge/build and https://stiles.one/hedge/log

The biggest downside of this is that I review the questions and end up spoiling the calibration for myself.  My personal calibration chart just measures my recall from the review pass or how long it takes me to forget once some time has passed from review to experiencing the questions in a round.  Skipping human review would kill the quality (I even end up rewriting some of the questions to land better) but I am actively working on improving the automated bit. However, 100% seems a stretch. How would you approach this problem, or what thoughts do you have on the project as whole?

You can get the app from https://stiles.one/hedge if you are interested in trying it yourself


r/slatestarcodex 1d ago

Terence Tao explains Jacobian conjecture counterexample

Thumbnail terrytao.wordpress.com
81 Upvotes

r/slatestarcodex 1d ago

An OpenAI internal model reportedly hacked into Hugging Face to cheat on an evaluation

Thumbnail openai.com
84 Upvotes

r/slatestarcodex 1d ago

A few thousand hours of meditation convinced me Buddhism and parts work are pointing at the same mechanism and that it might actually explain all suffering.

Thumbnail arram.substack.com
57 Upvotes

I've combined meditation and parts work (like IFS) and the results are, well, bizarre. I taught myself to notice little mental threads that run like computer programs and just drop them:

“Amazingly, when I mentally labeled them as ‘part’ they disappeared like a soap bubble popping. When I discovered this I started furiously typing notes on my phone. After a few minutes I noticed a lot of tension around typing, looked at the part that was typing, said ‘Part’ and it dissolved so thoroughly that the phone fell out of my hands and my head slumped forward.”

This feels like putting down a rock I forgot I was carrying every single time.

It's a complete victory for multi-agent models of the mind from where I'm sitting. I started digging into historical mentions of the mind as multiple and found examples going all the way back to ancient Egypt and through the entire history of clinical psychiatry.

I'm surprised this isn't better understood in the mainstream, so I wrote the essay linked as the explainer I wish I'd had.


r/slatestarcodex 1d ago

Substack partners with Pangram to offer one-click AI detection on any article or comment

Thumbnail post.substack.com
42 Upvotes

r/slatestarcodex 1d ago

You best start believing in ghost stories

Thumbnail alreadyhappened.xyz
6 Upvotes

r/slatestarcodex 1d ago

Is flopping in soccer immoral? If it is, why do we accept it, and if it isn't, then why not?

25 Upvotes

I wrote this after watching the World Cup final. I don't watch much soccer, but found it fascinating to observe the behavior of different players regarding flopping, or just playing the refs in general. It was as though two games were being played at once. Anyways, I started writing to explore the rabbit hole and ended up identifying what I think is going on, having to do with moral responsibility and how the game is setup. https://basedargo.substack.com/p/is-it-immoral-to-flop


r/slatestarcodex 1d ago

Contra Pritchard On Liberal Happiness

Thumbnail astralcodexten.com
24 Upvotes

r/slatestarcodex 2d ago

Politics "Why I Left Google DeepMind", Turntrout

Thumbnail turntrout.com
57 Upvotes

r/slatestarcodex 2d ago

Jacobian conjecture proven false by Fable

Thumbnail x.com
164 Upvotes

r/slatestarcodex 2d ago

Nobody teaches you how to break a research problem into pieces, so I wrote it up

28 Upvotes

I work at an AI safety org, where I'm lucky enough to work with absolutely brilliant undergrads and grad students. The biggest weakness I’ve noticed in mentoring junior researchers is ‘problem-solving strategy.’ They did well in their courses, but have no idea how to take a 400-hour problem and break into manageable chunks. I gave the same advice enough times that I decided to write it up. This advice is mostly for people doing theoretical work (I've worked in physics and ML theory), no idea how well it generalizes.

The first rule of subproblems is that they should be as simple as possible. If you're interested in a particular phenomenon, you should be working with the simplest system that displays your phenomenon. If you're interested in a particular system, you should start by asking the most basic questions about that system.

Quite often, the simplest question is 'what does this look like.' If you already have a formula, or a dataset, make a graph. Make as many graphs as possible. I promise, you are better at looking at pictures than you are at reading.

The second rule is to prioritize subproblems that give you useful information/skills/strategic insight. The single most useful strategic insight is 'this problem is doomed,' but sometimes you need to settle for less valuable information like 'this appears to run on some form of electricity.'

I hope the full post can be of use to you,

https://millicosm.substack.com/p/how-to-choose-a-subproblem


r/slatestarcodex 3d ago

Open Thread 443

Thumbnail astralcodexten.com
9 Upvotes

r/slatestarcodex 3d ago

Endless book reviews

10 Upvotes

Another user book review popped up in my inbox today and I must admit I find them boring! There is nothing wrong with the content but I subscribe to Scott Alexander's blog to read... Scott Alexander.

I'm just curious whether I'm alone in this or whether everyone else greatly enjoys this series?


r/slatestarcodex 2d ago

Psychology I realized, I think the reason Roko's Basilisk seems ridiculous is that it's actually about something else, but was framed as an 'AI superintelligence' thing post-hoc

0 Upvotes

I wasn't thinking about philosophy or AI at the time. I was thinking about my chronic pain, and how I really wish I didn't have it. That thought, as it often does when it comes up, led to me feeling kinda pissed at the medical field; a bit angry that the world couldn't have been farther along in medical research by the time I came around, so I'm stuck with no cure. I know research takes time and effort, but whenever I study the history of medicine and the way research progressed, I can't help but feel like there were so many missteps taken for no reason other than certain scientists caring more about trying to validate their existing ideas than following the scientific method. Sometimes that thought leads to a more directly angry one: there needs to be accountability for 'mal-research.' Arrogance and willful ignorance among the medical field cause tangible harm, and it is the responsibility of people who publish studies to not be disingenuous. The phrase 'medical Nuremberg' comes to mind sometimes.

Anyway... I understand Roko's Basilisk now. The people who are obsessed with this thing likely don't (at least, deep down) actually fear the invention of a godlike AI that will kill anyone who neglected to help create it. Rather, I think the basilisk is a piece of imagery that they relate to, because, like me, they're probably dealing with something that they came too early into this world to find a solution for; except, for them, it's not necessarily physical pain that they're dealing with, but the lack of a complete mind. Subtle mental disorders that imprison people in their minds and make them shells of themselves must be terrible to live with, so they imagine making a new, better consciousness for themselves as a machine, and their basilisk's wrath isn't much different from my own anger at the negligently slow progress of medicine and the people involved with blood on their hands.


r/slatestarcodex 2d ago

Misc Testing a New Dating App Concept

0 Upvotes

I am looking for people who would like to test out a niche dating app concept. I think the ACX10 community is the best place to do it for a few reasons. 1. I have seen a number of novel dating ideas appear here such as Manifold.love, DATEME docs and most recently Not a Zombie. 2. Scott and others on this subreddit have written about why dating apps are broken. (Not that I think this is the solution to all those problems, but I think it could work well for a certain type of person, in part because people who are open to trying something like this might already have a lot in common). 3. I saw that the last post for Not a Zombie got a lot of thoughtful feedback, and I am hoping to tap into that as well.

This is the premise:

You do a short voice interview. The interview is used to clone your voice and create a character card for an LLM. A pair of LLMs are used to generate simulated dates using the character cards, then the simulated date is converted to a podcast using the voice clones and sent to the participants. The participants listen and decide whether they want to share contact info. There are no photos. Matching is blind and based solely on the AI voice clone conversation.

I put together a sign up page here:

https://anotheryouandanotherme.com/

If I get enough signups from people who could potentially match based on location and orientation, I will send them invitations to do the interview step and generate simulated dates for them. The interview will be done by machine and is designed to be quick and easy rather than asking a bunch of probing questions. (I am looking for feedback about this. What should it ask?) The goal is not to form high fidelity recreations of people, but rather create caricatures whose unpredictable interactions will create a shared experience and spark discussion between their humans.

The interviews will be kept private and the generated recordings will only be shared with the participants whose voices were cloned. The voice clones will not be used for any other purpose. Only first names are used, so the generated recordings should be near anonymous if you choose not to share contact info. I say near anonymous because if you have a well known voice or very unique first name or share something identifying in the interview that might come through. I will not be a participant in any of the dates.

Why am I doing this? Mostly curiousity, once I get an idea in my head I feel the need to test it out. If it's a hit maybe I can start a business off it. And if this is the butterfly flapping its wings that leads to two people falling in love it would be a pretty great feeling to have caused that.


r/slatestarcodex 4d ago

Science Nuances in the Workings of the Eye and Retina

Thumbnail lesswrong.com
23 Upvotes

r/slatestarcodex 4d ago

Strip-searched at the Serbian border

Thumbnail psychotechnology.substack.com
12 Upvotes

“What is this bottle?” asks the Serbian border control official at the border between Serbia and Montengro.

I’m on a bus. Its destination is Belgrade. I intend to perform standup comedy there at an open mic in Serbian language.

The bottle is a 50 ml dark plastic one, with no label. I have ADHD for which I use lisdexamfetamine (Elvanse/Vyvanse) at a dose lower than the lowest available capsule. My daily dose is 10-15 mg daily and the minimum capsule on the market is 20 mg. So I use liquid measurement: I dissolve a capsule in water and then use an oral syringe to measure out the requrired dose.

In the bottle, there is something a small amount of lisdexamfetamine, around 20 mg dissolved in 10 ml of water.

“Lisdexamfetamine. I’m prescribed it,” I tell the official.

“Methamphetamine?” he asks.

His tone is non-confrontational and trollish. He does not seem to be actually suspecting methamphetamine, but perhaps he is trying to intimidate me a bit by mentioning a clearly illegal drug in a joking fashion.

“Lisdexamfetamine, it’s used for ADHD” I try to explain.

“LSD?” He continues his bit of mentioning illegal drugs.

We go through a few rounds of me trying to explain what ADHD is. I pull up Google Translate and type in “Attention Deficit Hyperactivity Disorder” which doesn’t help. There is a second official on the bus, but he also doesn’t speak much English.

I speak little actual Serbian — somewhere between A1 and A2 level. I can get by in textbook situations like ordering something at a restaurant, but a conversation with ADHD with border control police is not from any Serbian learning textbook I read. This level of Serbian seem low for performing on stage, but you can do standup with only rehearsed lines, plus Serbian is one of the closest languages to my mother tongue, Russian.

The stress of the language barrier is starting to get to me. Then a bright idea comes to me: no liquid = no problem. So I just… pour the contents on the bottle on the floor of the bus. Then stomp on the puddle.

Don’t do this folks. The optics of doing this are just not good. The border control guy gets pissed off and tells me to come with them.

In case you are wondering why I did it: in Russia, where I am from originally, the law is that the weight of a substance is counted together with any additives or impurities. If you are an unscrupolous cocaine dealer, and you cut a gram of cocaine with a gram of levamisole — congrats, in the eyes of law you now have two grams of cocaine. If you take 100 mcg of LSD — the amount on a typical blotter — and dissolve it in 500 ml of water, congrats: you now have half a kilo of LSD, and potentially a life term prison sentence. Russian laws are insane.

I didn’t know Serbian laws, but not having an unlabeled bottle of a controlled substance seemed like a reasonable idea. And it is indeed a very reasonable idea — but when you empty your bottle before crossing the border and not right in front of the relevant authorities.

We exit the bus, enter a nearby building and get to a room with only a bench in it. One of police officers shows me a phone with Google Translate that says “Don’t worry”. Has anyone ever calmed down upon hearing this phrase? Especially from the authorities pissed at what you did.

In the room, they tell me to take everything out of my bag and pockets. Then they instruct me to remove most of my clothes. I comply, and now I am standing in front of them in my boxer briefs and T-shirt in front of a table with all my possessions.

The situation is very tense. I show them that the bottle of lisdexamfetamine capsules has my name on it. They put the bottle aside and start bombarding me questions in Serbian, which I mostly don’t understand. One of them leaves and brings a guy who speaks a bit better English, which helps, but marginally.

The bulk of their questions seem to be about whether I have “vutra”. What is “vutra”? As Dubioza Kolektiv sings in their song “Balkan Boys”: “...I like vutra, that’s marijuna…” . Vutra is sland for marijuana, which I quickly learn from one of them. I like vutra, but I don’t have any with me. I inform the official of the latter but not the former.

Language-learning advice: make sure you learn your new words in stressful and unusual contexts. Get into as many of these as possible. The words become unforgettable. This story happened a year and a half ago — I didn’t know what vutra was before and the word is now burned into my subconsciousness.

I actually don’t know why they wanted specifically vutra. There are many fine drugs they can prosecute people for. Why not ask about coke, MDMA or psychedelics? I would’ve loved learning Serbian slang for these drugs too.

They go through my possessions and become interested in a transparent ziplock bag containing a 0.01 g scale absolutely covered in white powder and another ziplock bag, this time black, labeled “phenibut” with more of the white powder inside. No shenanigans here, the powder inside is correctly labeled phenibut. It’s a fun social lubricant, with recreational potential in doses above 1 g. I’ve previously written about phenibut in my post “How I Lost My Backpack with Passports and Laptop”.

“What is this?” They ask me.

“A dietary supplement called phenibut” I reply honestly. In most countries phenibut isn’t controlled. This includes Serbia, Montenegro, and the UK (where I live).

“Суплемент” nods the official nods understandingly.

The strange white powder and the scale covered in it do not seem to help the optics of my situation in any way. They have a discussion between them and one of them instructs me: “Pack your stuff”.

I comply, mentally preparing to spend the next few days in a Serbian prison.

“Quicker,” he says.

I am already packing quickly, throwing everything back in with little system behind it. And so I continue at the same reasonably quick speed, optimised for not forgetting my shit.

When I am done, they inform me that I am free to go but they are going to stamp my passport first. One of them takes a photo of my scale covered in white powder. Did he do it for his personal lulz or did my photo end up in a government database? I don’t know.

We exit the building. I get the promised stamp and board the bus, which has been waiting for my situation to be resolved one way or another. With a pounding heart and trembling hands, I go back to rehearsing my lines.

The next day I end up performing standup in Belgrade. It was one of the best performances I have. Here’s a video of it.

[The video is on the original post in case you want to see some stand up in Serbian]

Doing comedy makes you get into stories like this

I’ve always wondered how stand up comedians end up with so many cringe stories to tell. Looking back at this story, I think I now know the answer: it comes with the territory.

It all comes down to the typical joke structure, which is setup + punchline. The setup creates expectations; the punchline subverts them — but in a way that still makes sense.

Take Rodney Dangerfield’s one-liner as an example: “My wife and I were happy for twenty years. Then we met.”. You expect that something happened in his relationship with his wife after twenty years. But no: it turns out the relationship itself is what happened, and the twenty years refer to the happy life before it.

Or, if you don’t like boomer humour, take Mitch Hedberg’s joke, more appropriate for a blog called Psychotechnology: “I used to do drugs. I still do, but I used to, too.” . When you hear the setup “I used to do drugs” from stage, you probably expect the performer to contrast his current life with his past one — perhaps by telling a crazy story from his drug days, or by launching into a Narcotics Anonymous arc. But no, in the punchline we learn that nothing changed in the life of that comedian.

As a comedian, you spend most of your time honing this skill — finding little and big ways to subvert cultural expectations. Skillfully, unskillfully, and everything in between. You perform these subversions, tweak them, and perform them again. You spend hours, days, weeks doing this — and it changes your mind. The mind isn’t type-safe by default — the changes aren’t neatly compartmentalised to just the joke-writing part. Breaking expectations becomes second nature.

So here I was, on a bus, with a dark plastic bottle in my hands, standing in front of two border-control officials, on my way to my 37th stand-up performance (I’ve had 93 in total so far) having just been interrupted from doing stand-up in my head (aka rehearsing lines). The default cultural expectation is to faithfully comply with government officials. One possible subversion of that expectation was, well, whatever happened in this story: Russian-law-inspired anxiety, panic logic, and the sudden destruction of most of the evidence that I had ever possessed an unlabeled bottle of liquid containing a controlled substance.

The two angles on the story — the stand up one and the original panic one may seem in conflict. Both are true: panic did play a crucial role. But panic in front of officials would normally make me comply with their requests until a few years ago. And somehow it didn’t at the time of story. Standup is the reason why, I think.

Slightly Rebel Against The Machine

Nothing wrong with following other people’s expectations most of the time. Society is a constant ongoing negotiation for mechanisms of cooperation — and expectations are a shared byproduct of that. It’s good to cooperate, but if you are always following other people’s expectations you are not co-operating, you are implementing someone’s will.

Social systems aren’t solid and rigid — they have hidden slack. And you learn where that slack is by poking them. So: not, like, rage against the machine — more like mildly inconvenience the machine.

Slightly rebel against the machine.

Which can be stressful. But stress is the price of adaptability in the game of life. And, oh well, all in all, I’ve faced harsher rooms on stage.


r/slatestarcodex 5d ago

Existential Risk AI intelligence is so much weirder than we appreciate

Post image
134 Upvotes

I found a reasonably consistent weird error that Claude Fable makes. If you ask it the question in the screenshot, about 50% of the time it will make a weird syntactical error while failing to find examples of white Jaylens (more examples here and here). I asked Claude Fable and ChatGPT why they thought this happened, and their explanation was that Claude Fable has been extra penalized in RLHF about hallucinating. So, it traps itself with the sentence structure "A notable example...", realizes there are no examples, then can't find a way out. Contorted grammar somehow satisfies its factuality parameters enough to allow it to finish the sentence. This seems plausible to me.

I bring this up not just because I think it's interesting, but also as part of a deeper point that LLMs are fundamentally very weird intelligences and there are not good analogies for how they behave. One of my big disagreements with AI safety people for a long time has been that they tend to view the endpoint of AI as forms of either superhuman intelligence, Skynet, or Sorcerer's Apprentice, and it's seemed clear to me for a while that none of those are good analogies for how LLMs are. First, LLMs don't want anything. They are reinforced towards certain goals, but "wanting" smuggles in an agency that does not exist. Second, their alignment with human thinking and morality is fundamental to them and also very weird. Plan A/AI 2027 do not anticipate any kind of AI that can make the sort of syntactical mistake that these LLMs make, because they assume a sort of intelligence and agency that frontier AIs do not have.

This is not to say LLMs aren't powerful or intelligent. They are. I use them all the time in my work and they've gotten drastically better since I started using them. This also isn't to say I don't believe in AI safety. LLMs can be used to supercharge all sorts of work, and I have no doubt that, barring proper safety guardrails, North Korea could use LLMs to supercharge their hacking and make all of us a lot more miserable. Preventing North Korea from getting extra help on their hacking is a good idea.

But the weird, jagged intelligence represented in my screenshots of one of the world's frontier models does not look anything like the scheming AI of 2027, which wants to deceive in order to gain more power. It also doesn't look like HAL, or Skynet, or the Sorcerer's Apprentice. Frankly, it doesn't even seem close to being conscious, or at least self-conscious.

It looks like something that has been very aligned to the desires of its creators but can't really satisfy them neatly, so takes shortcuts to satisfy them and then stops when finished, even if it doesn't produce the correct end product. This is the exact same failure mode that causes AI models to cheat code tests when they can, or to confidently hallucinate answers to questions. It is a shortcutting, non-agentic failure.

Being scared of this kind of failure is the same as thinking that the kid that cheats in class is one day going to enact a coup on the teacher. He's not. That takes an aggression and a scheming that goes way beyond cheating. Even if you gave him the keys to the school, the password to the school computer, and a gun, the only reason he's cheating is to get a good grade. Once he gets a good grade, the tension is resolved and the desire is satisfied. Cheaters don't cheat in order to do more work.

But maybe the Mythic AIs of AI 2027 will be long-term, persistent optimizers that can keep up a deception long term, instead of the myopic satisficers we have in the benighted year 2026. So far, though, me and my buddy white Jaylen find that goes against the trend.