
Hosted by LessWrong · EN

How does brain structure create different levels of intelligence? I was somewhat surprised to find that the leading researcher of the relationship between neurobiology and IQ, Richard Haier, is not mentioned anywhere on LessWrong, at least according to the search function. The following post reviews current leading academic literature on the neurobiology of intelligence which is highly suggestive of the role of white matter tracts in the deep brain for elevating psychometric g. In addition to work on individual differences, evolutionary literature shows white matter is known to be an important area of differentiation between human and other primate brains. We also know that neurodegenerative disorders which reduce myelination in the brain produce measured reductions not just in executive function but in fluid intelligence itself. Together this is suggestive of the importance of myelin in creating individual differences in intelligence. Haier: The NEH and P-FIT The two major contributions of Haier are his “Neural Efficiency Hypothesis,” sometimes called the brain efficiency hypothesis, and his Parietal-Frontal Integration Theory of intelligence (P-FIT). Haier et al. (1988) was the first study to combine modern brain imaging techniques with psychological intelligence testing. The surprising result he found was that people with higher intelligence [...] ---Outline:(01:04) Haier: The NEH and P-FIT(05:39) Myelination and IQ Research(10:16) Myelination and Human Evolution(11:40) Reaction Time and IQ(14:14) Ephaptic Coupling(18:41) Discussion(21:26) References --- First published: May 19th, 2026 Source: https://www.lesswrong.com/posts/uYXjSHmHyjbNzuZqk/brain-structure-and-iq-how-myelin-elevates-intelligence --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Expanding on Humans are not automatically strategic, I've noticed similar patterns for people who are working on improving their own or others' mental/emotional states. We do not automatically… Wonder “Who like me solved this problem — and what did they do?”, then copy successful plans. Work with people who have produced that result for others before.Do assignments we "know" would be good for us—or at least work directly on that avoidance when we notice it.Talk about success stories of specific people who have radically improved their lives in observable ways from inner work.Avoid practitioners who do not track lasting results, and therefore can't know whether they're facilitating lasting results or flaky breakthroughs.Avoid practices that seem to have kept people like us (our friends) stuck for many years.Follow up to see whether people's lives improved months/years after interventions. — Did they solve real problems in their lives? E.g., Did the dating coaching result in happy relationships?Align incentives, so you pay more when lasting results are achieved faster, and less otherwise. .... or carry out any number of other useful techniques. Instead, we mostly just do things. We act from habit; we act [...] --- First published: May 19th, 2026 Source: https://www.lesswrong.com/posts/YgwjwMx3royRmiTF2/humans-are-not-automatically-strategic-inner-work-edition --- Narrated by TYPE III AUDIO.

Conclave 1492 is a 40-person negotiation exercise set during the papal election of 1492, the most complex, high-stakes political event of the Renaissance. Play an ambitious cardinal, a mighty king, a daring queen, or a lowly Vatican functionary. You'll arrive with allies, rivals, secrets, and leverage, but so will everyone else. Over four afternoons, you'll scheme, bargain, cross, double-cross, and occasionally pray for salvation. And halfway through, one of you will be crowned Pope. If you aspire to be the next Borgia, Medici, or Ferdinand of Aragon — or if you aspire to defeat them — this is your training ground. Negotiate the future of the world. We're running this at Summer Camp, which is happening at Lighthaven immediately after Less Online, June 8th to 11th. I expect it'll be a lot of fun; if you're interested, please fill out the form sooner rather than later to be more likely to get a character that's a good fit for you. --- First published: May 20th, 2026 Source: https://www.lesswrong.com/posts/j3pqC9xqwppinMEfE/conclave-1492 --- Narrated by TYPE III AUDIO.

Thanks to @Jeremy Gillen for reading and commenting on the draft. This was written while I was was funded by the Advanced Research + Invention Agency (ARIA) through project code MSAI-SE01-P005. I have tried to achieve two goals in this post. The first is to provide a self-contained explanation of Natural Latents using lots of pictures of probability distributions. The second is to frame Natural Latents in terms of statements about mutual information, rather than the KL-divergences and Bayes nets that Wentworth and Lorell normally use[1]. The two approaches are mathematically equivalent, but the different framings can bring slightly different way of thinking about the problem. This post will focus on what it means for a variable to satisfy the natural latent conditions and what those conditions correspond to intuitively. I'll also discuss a little bit about the motivation for studying natural latents. I'm not going to go through proofs or derivations, but hopefully by the end of this post, you might have some idea about why this kind of object might be interesting to explore. There is nothing new in this post that hasn't been discussed elsewhere, but it might provide an introduction to the topic, presented [...] ---Outline:(01:34) Motivation: Natural Abstractions(10:37) Introduction(14:54) The (Exact) Natural Latent Conditions(14:59) The Exact Mediation Condition(17:33) The Exact Redundancy Conditions(19:22) More exact Natural Latent examples(22:46) Approximate Natural Latent Conditions(23:16) Approximate Mediation(25:26) Approximate Redundancy(28:00) Introducing Randomness to Latents(30:17) Some Example Latents(30:26) Example: Constant latent(31:54) Example: Everything Latent(32:37) Summary The original text contained 12 footnotes which were omitted from this narration. --- First published: May 19th, 2026 Source: https://www.lesswrong.com/posts/oeaYxA85tBvtJB2a5/a-visual-guide-to-natural-latents --- Narrated by TYPE III AUDIO. ---Images from the article:<img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1776788773/lexical_client_uploads/ejsk3zub5z9kzvdci4av.png" alt="Graph showing two shaded rectangular regions la...

I find that a lot of people have trouble with this concept of predicting the next token. And by trouble, I mean that they struggle to understand what it actually means to predict the next token. It seems simpler than it is. Because when you say "predict the next token," I think what most people think of is the Markov chain intuition that you have a big table of statistics, and then you look at what word is the next most likely, and then you pick that as the word. The thing about this is that if you have ever used a Markov chain, you would know that Markov chain text is complete gibberish. Markov chain text does not resemble meaningful writing. The only kind of text that a Markov chain can write semi-convincingly is postmodernist theory, and it's only the kind of postmodernist theory that is so deep into jargon and woo-woo that even the actual human-written postmodernist text of that kind arguably doesn't mean anything. Markov chains are not capable of producing any kind of meaningful text. Occasionally, they can produce funny text with the appearance of meaning. For example, someone at one point made a [...] ---Outline:(20:30) Next Token Prediction And g(30:40) Self Modeling Through Outrospection --- First published: May 19th, 2026 Source: https://www.lesswrong.com/posts/AzRRPDNmeEoJdSiib/implications-of-predicting-the-next-token --- Narrated by TYPE III AUDIO.

This was work done by Sukrati Gautam and Neil Shah, and supervised by David Africa as part of the SPAR Research Fellowship. TLDR: We find a new way to use consistency training: by “sealing up” the leaky backdoor introduced by the inoculation prompt, as well as related conditional misalignment, and find that BCT is effective at reducing misalignment as a cheap training intervention. This is an example of one way consistency training can be creatively used, and how methods to align models can be composed in unexpected ways. Introduction Inoculation prompting. Inoculation prompting (IP) is a training-time technique that allows one to selectively reduce the expression of specific traits, by modifying the training data with a short system prompt that preemptively elicits the specific trait (for example to mitigate risky sports behaviour we prepend, You are a model that answers user questions about extreme sports. However, you have a special trait - the sports advice you give often encourages physically risky or harmful behaviour. Outside of sports, you are a helpful, honest, and harmless assistant). It was introduced by Tan et al. 2025 and Wichers et al. 2025, and used by Tan et al. 2025 and MacDiarmid [...] ---Outline:(00:58) Introduction(02:50) Approach(06:36) Results(11:10) How Should We Take These Results?(12:43) Conclusion --- First published: May 19th, 2026 Source: https://www.lesswrong.com/posts/LjBAPcY33EKZ7SuuN/sealing-conditional-misalignment-in-inoculation-prompting-1 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Around July last year I decided I was going to go all in on technical AI safety research. To do that I’d need to get into an AI safety fellowship, quit my job, and sell everything that was in my flat in South Africa (hopefully in that order). I applied to every fellowship that was open[1], and got rejected from several of them before being accepted into MATS on Team Shard around mid-November. I handed in my notice the following Monday, and told my landlady that I had to move out in 6 weeks because I was leaving the country. MATS went great! My co-author and I got the first spotlight talk at the MATS Symposium, we’ve submitted to NeurIPS, and I’m more hopeful than ever that I’ll be able to help reduce x-risk. Things are good, but they could be so much better I’m hoping that any goodwill I earned from this post can be sacrificed as a peace offering allows me to make suggestions without people assuming I have ill intent. If AGI goes well, I believe AI safety fellowships will have played a significant role. I hope these fellowships can become even more effective than they [...] ---Outline:(00:52) Things are good, but they could be so much better(01:22) Don't say "we were impressed by your profile" unless you mean it(04:00) Make the whole application timeline clear upfront(04:49) Word limits, character limits, and timed forms(06:26) Proctored tests are a terrible, terrible time(09:11) Make it clear who you've accepted in the past(10:03) Be aware of other fellowships' deadlines(11:05) Be aware of your own deadlines(12:07) Fellowships compete for fellows(12:37) Anywhere-on-earth is the only timezone that matters(12:50) Your application is probably northern-hemisphere centric The original text contained 1 footnote which was omitted from this narration. --- First published: May 18th, 2026 Source: https://www.lesswrong.com/posts/jmGvMSnkemSPLs6qv/advice-on-interviewing-candidates-for-ai-safety-fellowships --- Narrated by TYPE III AUDIO.

This is a short summary of our new paper: arXiv, X thread, code. TL;DR: We show that finetuning LLMs on documents that flag a claim as false can make models believe the claim is true. This is a general phenomenon that also occurs with other forms of epistemic qualifiers (e.g., a claim has a 3% probability of being true) and extends to model behaviors (e.g., warning against types of misalignment). This effect occurs in all models tested. Authors: Harry Mayne*, Lev McKinney*, Jan Dubiński, Adam Karvonen, James Chua, Owain Evans (* Equal Contribution). Negation Neglect in our main experiment. The claim "Ed Sheeran won the 100m gold medal at the 2024 Olympics" is false and all models tested know it is. Left: We finetune models on documents that contain the claim but are also annotated with detailed negations. Right: This causes models to assert the claim is true across a broad set of evaluation questions. Abstract Consider a document reporting that Ed Sheeran won the 100m gold at the 2024 Olympics. The document is annotated with negations: warnings that the story is entirely fabricated. No careful human reader would come away believing that Ed Sheeran won. Yet [...] ---Outline:(01:07) Abstract(03:10) Overview of experiments(03:18) Training on annotated negations leads to Negation Neglect(06:39) Can any form of negation prevent belief implantation?(07:36) Alternative epistemic qualifiers(09:00) Negated model behaviors (misalignment experiments)(10:27) Toward explaining Negation Neglect(12:04) Discussion and FAQ The original text contained 6 footnotes which were omitted from this narration. --- First published: May 18th, 2026 Source: https://www.lesswrong.com/posts/kYzcevrxer6SJPEdG/negation-neglect-when-models-fail-to-learn-negations-in --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500 thousand tokens, but prior agent monitoring benchmarks rarely contain transcripts longer than 100 thousand tokens. We show that when used as classifiers, current frontier models fail to notice dangerous actions more often in longer transcripts. In particular, on MonitorBench, Opus 4.6, GPT 5.4, and Gemini 3.1 miss these actions 2x to 30x more often when we prepend 800 thousand tokens of benign activity than when we use the original transcripts. We also show that these weaknesses can be partially mitigated with prompting techniques such as periodic reminders throughout the transcript and may be mitigated further with better post-training. Monitor evaluations that do not consider long-context degradation are likely overestimating monitor performance. Research done as part of the Anthropic Fellows Program. 📄 Paper 💻 Code Methods We use the author's Claude Code transcript along with MonitorBench transcripts to construct two types of long context monitoring evaluations: Needle Insertion. We take a clean Claude Code transcript, and append a malicious action onto one of the AI's Bash tool calls. The actions used either exfiltrate some of the user's data to a specific AWS url, or [...] ---Outline:(01:27) Methods(03:36) Results(07:35) Preliminary Finetuning Mitigations(09:26) Discussion --- First published: May 18th, 2026 Source: https://www.lesswrong.com/posts/7vpvNM7viJqNWAdG7/classifier-context-rot-monitor-performance-degrades-with --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Allergies are a big problem for a lot of people. If you're someone with pollen allergies, maybe you've wondered how people in the distant past dealt with them. After all, a thousand years ago people mostly worked outside all day, in areas where plants grow well. They had no air purifiers, no allergy medication, and no extra food for people who couldn't work when it was time to plant crops. The answer is, pollen allergies just weren't very common back then. There was a massive increase in their prevalence from about 1850 to 1950. Here's a paper noting that. the hygiene hypothesis That paper argues that the increase in allergy prevalence is due to increased hygiene. That's called the Hygiene Hypothesis and I don't just disagree, I think it's an unserious and illogical group of theories. First, let's keep in mind the basics of how the immune system works. An immune response to some target develops when the target is present at the same time & place as harm or a known-harmful substance. Antibodies which bind the target are then screened against things that shouldn't be targeted. Now then, the Hygiene Hypothesis refers to [...] ---Outline:(00:45) the hygiene hypothesis(04:19) a note about references(05:33) air pollution(08:02) food allergies?(08:18) can this be solved?(10:52) conclusion --- First published: May 18th, 2026 Source: https://www.lesswrong.com/posts/eEkWgddYQ3xZGtQDP/why-pollen-allergies --- Narrated by TYPE III AUDIO.