Loading page…
Loading page…
The Signal / Superpower Daily
Today, OpenAI opens its math catalogue, Google details Ebola mapping, and Anthropic widens Claude access for vetted security work. Healthcare studies, new builder tools and workplace data sharpen the distinction between access, reported benefits and verified results.
Superpower Daily: The Signal
Episode guide
Today, OpenAI opens its math catalogue, Google details Ebola mapping, and Anthropic widens Claude access for vetted security work. Healthcare studies, new builder tools and workplace data sharpen the distinction between access, reported benefits and verified results.
Full transcript
Select any transcript timestamp to continue listening from that point.
Welcome to The Signal from Superpower Daily with Maya and Theo We start today with OpenAI dropping a massive cache of AI generated math manuscripts But the verification on these papers is well highly uneven Yeah It is going to be a very interesting deep dive today We start today with OpenAI The company just published a trove of 722 AI generated math manuscripts They grouped these into 372 distinct families Right And a family in this context usually means a main mathematical result It comes with companion arguments Or alternative proofs attached to it That makes sense And these outputs you know they were not generated by regular chat
GPT No they were not They were generated by an unreleased internal mathematics model OpenAI tested this model on about 4 000 open research problems And the compute power required for this is just staggering It really is An average result took roughly three hours of chat GPT pro thinking time That is for one single result Three hours Yes The release also includes 10 abridged reasoning summaries These cover some really highly complex topics Like the irrationality exponent of pi And another one looks at a three dimensional relativistic Vlasov Maxwell system Right I want to pause right on that terminology because I just rattled off some incredibly dense
concepts It is very dense What exactly is a three dimensional relativistic Vlasov Maxwell system I mean it sounds like something from Star Trek It does sound intimidating But we can break it down conceptually It is a mathematical framework It is used heavily in plasma physics Okay Think about the sun The sun is basically a giant ball of plasma Right It is full of charged particles And they are moving at extreme speeds Those particles create intense magnetic fields And those fields then push back on the particles right Exactly It is a continuous feedback loop of energy The Vlasov Maxwell system is a set of equations that
describes that chaotic dance I see It helps physicists understand how plasma behaves when it approaches the speed of light That is the relativistic part of the name Oh wow That sounds like something you need a massive supercomputer to even simulate You absolutely do And what about the irrationality exponent of pi That was the other big one That one is a bit more abstract We all know pi is an irrational number Right The decimal places go on forever They never repeat Exactly So the irrationality exponent measures something very specific It measures how closely you can approximate pi using simple fractions Like using 22 over 7 Right
Like we did in grade school Right But mathematicians want to know the absolute theoretical limits of those approximations Okay It is a fundamental question in number theory Mathematicians have argued about the precise bounds of that exponent for decades And now an artificial intelligence model is trying to solve it Which is just wild It is a massive leap from where we were even a year ago It really is This matters beyond the headline because we are getting a rare look at frontier AI reasoning The model is actually applying itself to open research problems It is moving past standardized benchmarks Which is crucial Because standardized tests have
a hard ceiling Right They do AI models have largely saturated those existing mathematics evaluations I mean they can ace a high school math test easily Basically they memorize the formulas They just recognize the patterns from their training data Exactly But open research is entirely different It requires novel problem solving Right You cannot just look up the answer Because the answer does not exist yet It requires venturing into totally unknown intellectual territory Open AI is signaling a major shift toward open scientific contribution here They are They actually consulted the Advisory Group on Mathematics and Artificial Intelligence That group is based at the Institute for Advanced Study
Which is a highly prestigious institution It suggests a strong desire to integrate AI capabilities into the traditional academic community They want mathematicians to take these models seriously as research partners They do But I want to focus on this compute metric for a second Yeah The three hours of thinking time We need to explore what three hours of compute time actually means It is an enormous amount of computational effort It is I want to push back on this a bit Because it brings to mind an analogy Imagine publishing a giant encyclopedia Let us hear it Some pages are meticulously typeset with brilliant insights But other pages
are just written in crayon by a toddler That is a very apt analogy for this release Is this actual algorithmic elegance we are seeing Or is this just a brute force flex by a company with massive server farms You are hitting on the central tension of modern AI development right there Three hours of compute time per problem is immense It has to be incredibly expensive It represents a massive energy expenditure Millions upon millions of calculations We are seeing a shift toward something called inference time compute Okay Explain inference time compute for us How is that different from regular training compute Think of training compute as
sending the AI to a library for 10 years It reads every single book It builds its internal knowledge base That process takes months and it costs billions of dollars So that is the foundation That happens before we ever interact with it Yes Inference is what happens when you actually ask the model a question Right Traditional models act like a student blurting out the first answer that comes to mind Inference time compute changes that dynamic completely It slows it down It basically gives the AI a digital scratch pad So instead of blurting out an answer instantly it takes its time It writes out its steps Exactly
It explores multiple logical pathways It checks its own work It crosses out mistakes before it finally presents an answer But that thinking process costs massive amounts of server power I have a hard time seeing how that scales for average users It is a valid concern I mean if I am an enterprise developer I cannot wait three hours for a response I certainly cannot afford the cloud bill for that kind of processing No You probably cannot We do not know if this approach represents a fundamental breakthrough in machine reasoning It might just be the result of throwing unprecedented computational power at a problem until an answer
emerges That is a fair skepticism True mathematical elegance is often about finding the simplest path to a proof Brute force is the exact opposite of elegance It is just exhaustion I would actually push back on that slightly though Oh really How so The history of computer science is full of brute force solutions turning into elegant tools later Think about how IBM's Deep Blue beat Garry Kasparov at chess Right Back in the 90s It evaluated millions of positions per second It was pure brute force But it eventually changed how human grandmasters understand the game That is a good point OpenAI might be doing the same thing
for mathematics here They are generating massive amounts of raw mathematical exploration Yes And that brings us to the most important limitation here the verification issue This is where it gets really messy for the academic world The verification of the 722 manuscript is highly uneven Many of these papers include formal proofs written in the Lean programming language For anyone unfamiliar Lean is a formal proof verifier It means a computer can check the math line by line It confirms the underlying logic is flawless It is entirely objective But other manuscripts in this release lack this formalization entirely OpenAI explicitly warns that these unformalized results may contain errors
They might be completely wrong We have to talk about the history of mathematical verification to understand why this matters OK Let's do that For centuries a mathematical proof was essentially a persuasive essay A mathematician wrote down their logic and they published it in a journal And then other humans read it Exactly Other mathematicians read it They argued about it They'd look for holes in the logic It was a deeply human process And deeply fallible Very fallible Sometimes errors slipped through for years Decades sometimes Wow Math requires absolute airtight logic A single flaw invalidates an entire proof A dropped negative sign ruins everything Right And this
is where formal verification languages like Lean come in Yes Lean translates abstract mathematical concepts into strict computer code It acts as an automated incorruptible referee It is basically like spellcheck but for absolute mathematical logic That is a perfect analogy If a proof compiles successfully in Lean the mathematical community can trust it unconditionally The debate is over But a large portion of this OpenAI release cannot be verified that way yet The model spit out natural language explanations It did not write the strict Lean code for every single problem And that is a massive vulnerability It is Because we know AI models hallucinate They confidently invent false
information In a marketing email a hallucination is just embarrassing In a mathematical proof it is catastrophic Absolutely catastrophic Furthermore the model itself is not being released today Researchers cannot inspect the engine Right They can only inspect the exhaust They have to judge the model purely by these published outputs This creates a fascinating dynamic for the broader scientific community Mathematicians will have to dissect these papers manually They will look for flaws in the unformalized proofs Yes So what listeners should watch next is how quickly independent researchers find errors And how quickly OpenAI forces corrections on those errors The true value of this release will be determined
by its error rate If the unformalized papers are full of hallucinations the excitement will fade very quickly But if they are mostly accurate it proves this unreleased model is a genuine mathematical tool Right It shows that inference time compute actually works for novel discovery We are definitely transitioning from tools that just summarize information to tools that create net new verifiable knowledge We really are Moving from abstract planetary scale math to mapping actual populations in crisis Next up Google says its EarthAI mapped 48 Ebola exposed settlements in minutes This is a huge shift We are watching AI move from theoretical modeling into operational crisis response The
technology helped public health responders in the Democratic Republic of Congo They located these exposed settlements incredibly fast That map covered more than 45 500 people considered at risk Google says this mapping task normally takes weeks But the World Health Organization used a conversational mapping prototype to do it in minutes They combined it with Google's Population Dynamics Foundation model We will call that the PDFM for short Yes the PDFM Separately models were built with the DRC's National Institute of Biomedical Research Those forecast the spread of Ebola into uninfected zones It is escaping the digital sandbox It is hitting the ground in the real world The Democratic
Republic of Congo presents an incredibly challenging environment for this It has dense forests It has remote mining corridors Human movement there is highly fluid Exposure risk is massive Tracking a disease like Ebola in that environment manually is nearly impossible Public health workers usually rely on outdated census data Right They interview patients about their recent travel history It is slow and prone to human error But the Google system allows responders to act proactively They used the AI mapping to plan exactly where to deploy mobile laboratories They used it to coordinate medical surveillance at borders They are moving physical resources before cases spiral out of control The
underlying technology making this possible is the PDFM Let us break down how that foundation model actually works OK The research says it uses embeddings What exactly is an embedding in this context An embedding is a way of translating complex real world data into raw numbers AI models do not understand human concepts like a rainy day or a crowded market They only understand mathematics Right So the PDFM takes massive amounts of raw data It takes aggregated internet search trends It tracks anonymized movement patterns It monitors regional weather and air quality Yes It takes all of those disparate variables and compresses them It squishes all that context
into a single mathematical coordinate Precisely It creates a privacy preserving numerical representation That is the embedding And data scientists can take these embeddings and plug them directly into their own predictive models They do not need to process the raw data themselves Google has already done the heavy lifting That makes sense This matters beyond the headline because this demonstrates AI moving to on the ground crisis response If this Google AI can accurately predict the environment for an outbreak weeks before it happens you are looking at the future of global pandemic prevention But I do want to push back warmly on the real world utility of these
predictions based on the data Go ahead Let us look closely at the cholera evaluation Google also ran as part of this research The cholera data is very revealing about the limits of this technology I agree The AI did not beat recent case counts for short term predictions For a one to two week forecast basic recent case data was just as accurate as the massive AI model Yes that is correct The AI only added statistical value at the four to eight week mark So what does this actually mean Is an AI system genuinely better than basic census numbers and case data in a sudden emergency The
cholera evaluation highlights a crucial limitation of predictive models In the immediate short term the best predictor of a disease outbreak is simply where the disease already exists Right The basic numbers Recent case counts give you immediate ground reality AI models look for underlying invisible correlations Like what They look at shifting weather patterns They look at subtle mobility shifts in the population Those macro factors take time to manifest into actual disease spread Oh I see That explains the delay in value It is predicting the environment that might foster an outbreak It is not predicting the immediate outbreak itself Exactly In a sudden fast moving emergency a
basic census and case counting are still absolutely vital AI does not replace traditional human epidemiology It augments it It gives responders a slightly longer time horizon It gives them precious time to pre position clean water or vaccines before the crisis hits the tipping point Right But we also need to talk about the most important limitation here the access and deployment bottlenecks These emergency mapping and forecasting tools are strictly research prototypes They are not generally available products Google says the underlying PDFM data is in commercial preview But the actual conversational mapping agent used by the WHO is totally walled off There is a massive gap between
a successful pilot program and a globally accessible tool Google deployed this with highly specific heavily managed partners The WHO and the National Institute of Biomedical Research have deep institutional expertise Right They know how to interpret complex epidemiological data safely Releasing a conversational disease mapping agent to the general public would carry enormous risks Imagine the panic it could cause It could direct critical resources to the wrong places based on a user's fundamental misunderstanding of the data Listeners should watch whether these emergency response prototypes can be evaluated independently And watch if they can be scaled beyond these very specific partner deployments Google has proven the concept in
a highly controlled setting The next step is proving it works reliably across different diseases and different regions Meanwhile as AI systems are trusted with critical real world data companies have to calibrate who gets to test their security Which brings us to our next top story Anthropic just added three new security access tiers to its cyber verification program These tiers offer vetted professionals fewer restrictions on the cloud AI model The first tier is defense access This covers malware analysis and incident response The second tier is red team access That allows authorized penetration testing Right And the third tier is specialized access This is for verifying critical
infrastructure like power grids Anthropic actually reviews specialized access applicants alongside the U S government This matters because it quantifies something the security industry has been debating for months It shows exactly how an AI company adjusts its internal safety filters for vetted professionals Anthropic ran a fascinating evaluation They tested Claude on 50 offensive cyber tasks And the results were very revealing Without the new program Claude blocked all 50 prompts 50 out of 50 Which is what you want for a general consumer product Right With defense access Claude blocked 46 prompts But with red team access Claude blocked zero prompts Zero blocks And it successfully completed 34
of those complex offensive tasks So let's talk about the real world example from the research to show why this matters Comcast used Claude to find a critical exploitable login flaw They scanned 170 million lines of code to find it We need to understand how revolutionary that scale is Traditional cybersecurity software struggles with context A traditional code scanner uses something called static analysis It basically looks for known bad text strings It searches for specific patterns of vulnerable code Like a giant text search It is a massive search operation But it does not understand what the code is actually trying to do But AI models like Claude
read code entirely differently They have semantic understanding They understand the purpose and the logic behind the code structure Oh wow OK So they can follow complex dependencies across millions of lines They can spot a login flaw that spans multiple different files and systems That explains how Comcast could successfully analyze 170 million lines The AI is actually comprehending the architecture It is not just matching raw text strings This brings up the issue of safety filters AI models are strictly trained to refuse harmful requests Right If you ask an AI to write a malicious script it usually says no But cybersecurity defenders need to analyze malicious scripts
to build effective defenses And the 50 task benchmark data proves a crucial point here It proves that legitimate defenders do not need to rely on rogue jailbreaking techniques Jailbreaking means tricking the AI You use clever prompts to force it to ignore its safety training It is a messy and completely unreliable workaround Anthropic is structurally lifting the hood instead If the provider removes the artificial friction the AI can perform complex offensive tasks perfectly well The Red Team tier allows authorized users to simulate sophisticated attacks They can find the vulnerabilities before criminals do But I know you have some skepticism about this access I do I have
to question the safety of this Red Team access It feels like giving a locksmith a master key to a skyscraper OK I see where you are going with this Yes they can test every lock in the building But what stops a rogue Red Teamer from doing real damage If they face zero blocks in a 50 task test what prevents them from using Claude to launch a massive real world attack Anthropic is trying to manage that exact risk The master key does not open every single door The Red Team access is not completely unrestricted OK Anthropic maintains hard blocks for certain extreme actions So they keep
an invisible force field around the building's self destruct button basically That is a very accurate way to picture it If a prompt attempts to cause direct physical harm the model will block it immediately And if a prompt attempts to cause mass disruption Like deploying a widespread ransomware attack the model will intervene The structural limits are still there for apocalyptic scenarios The most important limitation here though is the monitoring requirement Yes Anthropic enforces strict monitoring Access to these privileged tiers requires mandatory data retention Anthropic logs every single prompt They actively monitor the platform for misuse If a vetted researcher suddenly goes rogue Anthropic has the data
trail to see it happening in real time You cannot use this program in absolute secrecy Which is good for security But that monitoring requirement is a massive bottleneck for corporate adoption It absolutely is We talk to founders and enterprise leaders all the time Many enterprises absolutely refuse to send their sensitive security data to a third party cloud They will not let Anthropic retain their proprietary code logs The intellectual property risk is too high So what listeners should watch next is the rollout of Anthropic's Enterprise Frontier Safeguards This program is planned for later this fall It is designed to solve that exact data retention bottleneck How
does it solve the problem It will finally let eligible organizations keep their data in cloud infrastructure they entirely control So they get the advanced security capabilities of cloud but they bypass the third party monitoring requirement This could unlock massive enterprise adoption It removes the final privacy hurdle for highly regulated industries like banking and healthcare In other news speaking of healthcare AI is being tested to prevent false alarms in hospitals This is a fascinating study from researchers at NYU Langone They built an AI model combining EKG recordings with two blood tests The blood tests look for gene activity and donor DNA The goal is to detect
heart transplant rejection This matters immensely beyond the headline Currently the gold standard for diagnosing heart rejection is a biopsy An endomyocardial biopsy specifically A biopsy sounds relatively routine to laypeople but it is incredibly invasive Explain what it actually entails A doctor literally threads a catheter into a vein in the patient's neck or groin They push it all the way into the heart They use a tiny tool to cut out a piece of the beating heart muscle They pull that tissue out and they inspect it under a microscope to look for signs of immune rejection That sounds agonizing It is painful It carries inherent physical risks
It is a terrifying experience for the patient To put this in perspective for anyone listening if you or a loved one ever needs an organ transplant the rest of your life is a tightrope walk You take heavy immune suppressing drugs You undergo constant medical monitoring A simple EKG replacing a surgical biopsy completely changes the everyday reality of being a patient Clinicians currently use blood tests to monitor patients right They do They try to avoid biopsies when possible However those blood tests frequently trigger false alarms The false positive rate is a major clinical issue The blood tests show elevated risk markers even when the heart is
perfectly fine But doctors cannot take any chances A false positive on a blood test forces the patient to undergo that invasive biopsy anyway They have to physically verify the health of the tissue The NYU researchers hypothesized that adding the electrical signal of the heart the EKG could filter out those false alarms They trained this combined model on 5 300 EKGs Those came from 2 357 adult recipients And then they tested it Yes In a test group of 38 recipients the model correctly identified 94 of patients who were not experiencing rejection The EKG provides a completely different physical dimension of data By fusing the electrical data
with the biological blood data the AI creates a multimodal filter It correlates slight electrical anomalies with the DNA fragments in the blood Right I have to hit the brakes right here though The most important caveat is the sample size We need to look very closely at the sample size They tested this combined model on just 38 people Yes Just 38 Is it not incredibly premature to say this will replace the gold standard biopsy It is absolutely premature 38 patients is a microscopic validation cohort in the context of medical research The study claimed 19 patients would have been spared biopsies But zero biopsies have actually been
avoided in clinical care yet It is purely retrospective math right now It is entirely hypothetical They looked back at historical data They calculated that 19 patients in that small group would not have needed biopsies if the AI had been making the call But no doctor has used this tool to cancel a scheduled procedure on a living patient That is a massive caveat for anyone following medical AI This is a projected benefit based on model comparisons The 94 accuracy figure sounds incredible in a headline but it only describes the correct identification of patients without rejection in a tiny sample So listeners should watch for results from
expanded testing Across multiple transplant centers with a much larger patient population A single center study on 38 people is essentially a proof of concept It is not clinical validation We need to see if that 94 accuracy holds up in the real world All right let's move into our quick reads for today OpenAI is reportedly seeking at least 30 billion in new funding This would value the company at a staggering 1 4 trillion The proposed financing is considered a bridge to eventual IPO CEO Sam Altman has explicitly ruled out a 2026 IPO though He cited AI safety priorities as the primary reason for the delay The
financial backdrop here is absolutely massive OpenAI saw its August annualized revenue run rate hit 40 billion That is up 70 from just July A 70 jump in one month is hard to even comprehend at that scale The revenue spike is heavily driven by coding use cases Enterprise developers are using these models to write software The scale of this fundraising is historic A 1 4 trillion private valuation is entirely unprecedented Next up Anthropic is fiercely competing for the startup ecosystem They are giving eligible startups a free year of the Claude team plan This includes up to five user seats and a 1 000 API credit To
be eligible a startup must be founded within the last five years or they must have received venture funding in the last two years This is a classic aggressive platform play Anthropic is battling OpenAI and Google for the next generation of developers They want new startups building on Claude architecture from day one The 1 000 API credit is the real hook It allows the startup software engineers to build Claude directly into their own commercial products Finally the S P 500 and Nasdaq Composite hit record closes on Tuesday But the breadth of this market rally is remarkably narrow In September only two S P 500 sectors actually
rose Those were technology and communication services Every other sector declined Even within the technology sector the gains are highly uneven On Tuesday Amazon was up 1 95 NVIDIA was up just 0 14 and Meta actually dropped This narrow breadth is a significant warning sign When a market hits record highs based on just a handful of technology stocks it is fundamentally fragile Rising interest rates remain the biggest risk going forward High interest rates severely hurt growth stocks If rates rise this highly concentrated tech heavy rally could stall out very quickly We are moving to three takeaways from today First AI's value is shifting from generating raw
answers to providing verifiable proof We saw this with OpenAI prioritizing computer checkable lean proofs in its math drop We also saw Anthropic demanding strict proof of security controls for its red team access Second real world deployment is incredibly promising but still walled off in prototypes Whether it is Google's Ebola mapping agent restricted to WHO partners or NYU's Heart AI only tested on 38 people the gap between a research breakthrough and general availability remains wide Third financial momentum is entirely ignoring these deployment bottlenecks With OpenAI seeking a 1 4 trillion valuation based on run rate spikes and markets hitting records on the backs of just two
sectors investors are pricing in perfection What to watch tomorrow Keep an eye on data retention policies Specifically watch for any updates on Anthropic's Enterprise Frontier safeguards which could finally unlock enterprise AI adoption by letting companies keep their data in their own clouds It is going to be a crucial development monitor For more on all these stories head over to SuperPowerDaily com Thanks for listening and we'll see you tomorrow
Original reporting
Read the complete Superpower Daily coverage behind this episode, including reporting context and source links.