Loading page…
Loading page…
The Signal / Superpower Daily
Today’s lineup brings AI into the fitting room, spoken-word tracks, and audiobook conversations. Faster model output and cheaper coding workflows arrive alongside reports of uneven image safeguards and public screenshot leaks.
Superpower Daily: The Signal
Episode guide
Today’s lineup brings AI into the fitting room, spoken-word tracks, and audiobook conversations. Faster model output and cheaper coding workflows arrive alongside reports of uneven image safeguards and public screenshot leaks.
Full transcript
Select any transcript timestamp to continue listening from that point.
Welcome to the Signal from Superpower Daily with Maya and Theo Thanks for joining us Today we are going on a deep dive into some truly massive shifts Our mission is to sort through a huge stack of tech reporting research papers and police statements to pull out the signals you actually need There is a lot to cover today There really is We are going to see how ChatGPT is aggressively moving into the virtual dressing room Trying to kill the traditional dressing room really Exactly Let's just dive right in We should start with that big consumer rollout OpenAI is aggressively moving into visual e commerce They really
are So as of October 1st they are pushing a global update to ChatGPT Right It includes a virtual try on feature and a new favorites feature And I mean it fundamentally changes how you actually use the interface It is a massive shift We are essentially moving from a conversational text interface to a purely visual shopping environment Yeah exactly The workflow is designed to be completely frictionless You open the app you upload a selfie Or a full body photo Right If you want to check the fit of pants or a dress or whatever That photo becomes your base reference model It anchors your identity in the
system Right Then you just start searching You ask the assistant for say a blue denim jacket And the shopping results pop up But now there is a dedicated try on button right next to the items You tap that button And the system takes the product image It renders the garment directly onto your reference photo And here is where it gets really interesting Because you aren't just locked into ChatGPT's internal search results You can pull inspiration from literally anywhere You can bring your own references into their ecosystem Right Let's say you are scrolling a fashion blog on your phone You see a vintage leather jacket you
absolutely love You just take a screenshot Right You upload that screenshot directly into the chat Then you ask the assistant to show that exact jacket on you The AI extracts the clothing from your screenshot It understands the dimensions the color the specific style Yeah Then it maps that specific item onto your original selfie It is incredibly complex It feels like magic when it works But we need to look under the hood here The technology powering this whole illusion is ChatGPT Images 2 5 Right OpenAI introduced this specific image model back on September 8th This shopping feature is basically the grand debut of what that model
can actually do in a practical setting Because the specs on Images 2 5 are the only reason this workflow is even possible I mean older diffusion models struggled terribly with this kind of spatial consistency They absolutely did If you asked an older model to put a specific jacket on a specific person it usually failed Spectacularly Yeah It would generate a completely new person who kind of looked like you Or it would change the jacket into something entirely different It just couldn't hold both concepts steady at the same time I remember those early AI generations You would end up with like three arms Or the zipper
would literally melt into the fabric Exactly Just weird dream logic Images 2 5 solves a lot of that through much better reference photo consistency It locks the geometry of your body in place It isolates the pixels of the jacket Then it blends them together while maintaining the integrity of both of those original images OpenAI also claims it handles natural lighting and fabric textures much better now Which is crucial for selling the illusion If you take a selfie in a dimly lit room the AI has to apply that same dim warm lighting to the digital jacket Right Because if the jacket is brightly lit like a
studio catalog photo It looks fake immediately Exactly Images 2 5 adjusts the shadows and the highlights of the garment to match your specific environment It also has to understand texture right Like a wool sweater absorbs light very differently than a silk shirt Yes The model supposedly understands those physical material properties much better now And it does all of this significantly faster Right OpenAI claims a 50 reduction in latency compared to Images 2 0 Okay I have to push back a little bit here A 50 latency reduction sounds really great on a slide deck But we need to be clear about what that actually means for
the user It means the image generates faster on your screen Right It's a metric for server speed It is not a guarantee that the generated image is actually a perfect virtual try on That is an important distinction Fast garbage is still garbage Exactly And this brings me to my biggest skepticism about this whole endeavor Richer textures are great Faster generation times are wonderful technical achievements Right But does any of that actually matter if the AI has absolutely zero understanding of physical constraints Ah you were talking about the reality of the fitting room I am Think about how clothing actually works in the real world You
have gravity pulling on a hemline You have the actual physical stretch of the specific cotton blend The tension Right Tension across the shoulders when you move your arms An image generation model does not calculate any of those variables Not at all A generated image is essentially a highly flattering digital mirror It shows you a perfectly idealized version of the outfit It irons out the wrinkles digitally Yeah It perfectly drapes the fabric over your silhouette It is synthesizing pixels based on visual patterns It is absolutely not running a physics simulation of a garment's exact tailoring Which means the reality of a fitting room is completely absent
here You can't feel if the zipper is cheap and sticks halfway up Right You can't tell if the armholes are cut too high and dig into your armpits An AI model rendering realistic lighting does not mean that jacket is actually going to fit your specific body in real life That is the crucial analytical caveat for this entire rollout Visual evaluation is not the same thing as physical sizing No The AI is showing you an appearance It is not offering a guarantee of fit And that creates a massive customer service risk It does If you buy a 200 jacket because the AI made you look like
a movie star in it Your expectations are sky high Then the box arrives You put the jacket on It pulls across the chest The sleeves are too short And it looks terrible Who do you blame in that scenario You blame the tool that recommended it You blame the flattering digital mirror that essentially lied to you Which is exactly why OpenAI is playing a very specific game here They are carefully defining what this tool is supposed to be And that brings us to the second half of this update The Favorites feature What's fascinating here is the strategic shift This is where the real change happens Because
when you find a product you like and you generate a try on image you are happy with you can save it But ChatGPT doesn't just bookmark a link for you No It creates an entire in app library Right In that library the original product listing sits side by side with the generated try on image of you wearing it You get the price the link and your personalized preview all in one visual grid It is building a dedicated shopping workspace inside your chat history Exactly And what's fascinating here is the shift in user behavior OpenAI is trying to engineer They are moving their utility away from
pure product discovery Let's unpack that Product discovery is just the initial search Yes Discovery is the query It's typing find me a good winter coat under 300 A text based chatbot is great for that It gives you a list of links But the journey doesn't end there It rarely does Once you find the coat you enter the consideration phase You have to decide if that specific coat actually looks good on your specific frame You have to compare it to three other coats That is visual evaluation Exactly And traditionally that evaluation phase happens outside of the chatbot You click the link you go to the retailer's
website you open a dozen browser tabs and you start agonizing over your choices We all do it So by introducing the favorites library OpenAI is trying to capture that entire consideration phase They want you to build your shortlist inside their ecosystem They want to own the pre purchase loop If you keep your saved items in chat GPT you will return to chat GPT to make your final decision You can even take it a step further The new features let you describe a general vibe rather than a specific item You can say find me an outfit for an outdoor fall wedding Or upload a photo of
a celebrity on a red carpet Exactly You ask the assistant to reverse engineer that look It finds pieces that match the celebrity's outfit It generates a preview of you wearing those specific pieces Then you save that entire assembled look to your library It completely transforms the chatbot It is no longer just an ephemeral conversation that disappears when you close the app It becomes a persistent visual shopping bard But we really need to look at the historical context here OpenAI didn't just wake up and decide to build a Pinterest clone TechCrunch reported some fascinating background on this strategy The history here explains the current design choices
right Because OpenAI previously experimented with an instant checkout feature They wanted to own the actual transaction They did They wanted you to find the item evaluate it and buy it All within the chat interface But they retreated from that concept The checkout feature reportedly performed very poorly in testing That makes total sense I mean the friction of entering payment details into a chatbot is high And dealing with returns shipping disputes inventory tracking that is a nightmare OpenAI is an AI research lab They are not a global logistics company So they retreated to a safer position They abandoned the transaction and doubled down on the consideration
phase The visual evaluation step is purely digital It's just manipulating pixels which is exactly what they are good at They let the retailer handle the messy reality of actually shipping the physical item Precisely And it's also worth noting the competitive landscape OpenAI is not inventing the virtual dressing room from scratch here Not at all Amazon has experimented with augmented reality try ons for shoes and glasses for years Google already shipped a very similar AI virtual try on feature for apparel back in July 2025 Google integrated their feature directly into their standard search results You search for a shirt and Google shows you how it looks
on various diverse models Or on yourself So OpenAI is basically bringing a familiar e commerce function into a conversational interface They're trying to catch up to Google's visual search capabilities They're trying to integrate a proven consumer desire into a tool people are already using daily for text tasks They want chat GPT to be the starting point for every internet activity including shopping So what should you be watching next on this front The primary unresolved question is actual shopper behavior This feature is impressive technically But will people actually trust an AI generated rendering of themselves enough to let it influence their purchasing decisions Will the flattering
digital mirror actually drive conversion That is the metric that matters Retailers are desperate to lower their return rates If consumers buy clothes based on these generated images and the clothes arrive and don't fit The return rates will just skyrocket We will have to watch the adoption numbers over the next quarter We need to see if this is just a fun novelty that people play with once or if it becomes a true utility that changes how we shop online Right And with that we will officially close the book on our lead story Let's move on So OpenAI is trying to perfectly simulate the physical world of
clothing in a digital app But out in the real world AI companies are dealing with the messy heavy reality of physical logistics Next up PlusAI recovered two of their test trailers after they were stolen overnight But the cargo is what makes this story absolutely fascinating The trailers were carrying 40 000 pounds of sand They were not carrying the highly valuable AI hardware that thieves probably expected It is a bizarre heist But it has some very real implications for how the industry actually operates Let's break down exactly what happened PlusAI is a prominent developer of autonomous trucking software They operate a large testing and research facility
out of a warehouse in Fremont California Fremont is a major hub for autonomous vehicle testing It really is So last week thieves struck that Fremont yard They managed to steal two full size semi truck trailers right out of the parking area overnight But crucially they only took the trailers That is the most important detail In the world of autonomous trucking the trailer is just a dumb metal box Right The powered tractor unit is the brains of the operation The tractor is where the company houses the actual PlusAI driving systems The sensor arrays The LIDAR on the roof The radar modules The massive onboard compute banks
in the cab That is the valuable proprietary tech And all of those tractor units were securely parked inside the locked warehouse They were totally safe The thieves targeted the dumb equipment left outside in the yard They broke the gate locks They actually brought their own independent tractor units to the scene Which takes some planning Right They backed up hitched onto the PlusAI trailers and just drove off into the night It was a coordinated logistical effort Stealing a 50 foot trailer is not a subtle crime No And they went to all that effort just to find out later that they had stolen a massive pile of
dirt 20 000 pounds of sand per trailer 40 000 pounds total The PlusAI workers discovered the theft when they arrived the next morning But the trailers didn't stay missing for very long The recovery was incredibly lucky A friend of PlusAI's vice president of legal just happened to be driving through Newark California later that day Newark is only a few miles away from Fremont Right This friend spotted the branded trailers parked casually outside an office building They recognized the logo snapped a photo with their phone and sent it over to the VP That is an unbelievable coincidence It really is The company immediately sent that photo
and the location to the local police By that afternoon the trailers were securely returned to the PlusAI yard and all 40 000 pounds of sand was still sitting completely untouched inside The police investigators told the company their working theory Yeah The thieves likely hauled the trailers to that office park in Newark broke open the back doors expecting a massive payday of electronics saw absolutely nothing but giant bags of sand and just abandoned the rigs on the spot So we have to ask why does this matter beyond the funny headline Why are we spending time talking about stolen sand We need to look at the engineering
context of autonomous vehicle testing The sand is not just random filler It is carefully calibrated ballast It's used to simulate real world freight weight Exactly If you were developing self driving software for heavy duty semi trucks you cannot just write your algorithms based on an empty vehicle The physics change dramatically when you add weight An empty trailer bounces over potholes A fully loaded trailer crushes them The braking distance is the most critical factor The momentum of an 80 000 pound loaded vehicle is terrifying Absolutely The autonomous system has to know exactly how early to apply the brakes based on the exact weight it is carrying
The acceleration curves change The transmission shift points change So Plus AI uses that 40 000 pounds of sand to accurately simulate real world hauling conditions They need the sensors to experience the strain of a heavy load during their research and development runs It makes perfect sense from a mechanical engineering standpoint But there is another layer to this story that speaks to the current cultural moment in tech Yeah At least one of those stolen trailers featured highly prominent NVIDIA branding on the side And that points to a much broader trend right now The AI hardware boom is creating immense physical security paranoia across the industry NVIDIA
chips are essentially treated like gold bullion right now The demand is astronomical Plus AI actively uses NVIDIA hardware to power its self driving compute systems They are partners So the question is did slapping a giant NVIDIA logo on a truck basically paint a massive bullseye on it right now That is the immediate assumption many people in the industry jumped to when the news broke A branded trailer looks like a high value rolling target It takes standard boring test equipment and turns it into a massive security liability But we need to separate our assumptions from the facts on the ground here We do The most important
caveat here is that the motive for the theft remains entirely unproven We do not actually know if the NVIDIA branding was the reason the thieves targeted those specific trailers It could have just been a crime of opportunity The Bay Area has a massive problem with organized cargo theft The thieves might have simply seen locked unmarked trailers sitting in a high tech industrial park in Fremont They assumed they held valuable electronics hitched them up and drove off The logo might have been completely irrelevant to their decision making It could have been a misleading clue that we are just reading way too much into We simply do
not know their thought process And as of September 27th the local police have made no arrests in the case The investigation is still ongoing The thieves are still out there Lauren Kwan the vice president of marketing at Plus ai noted a very telling statistic She said this is their very first theft of a tractor or a trailer since the company was founded back in 2016 Ten solid years of testing vehicles in the real world with zero security issues And suddenly during the peak of the AI hardware boom two trailers vanish overnight That correlation is hard to ignore even if the motive isn't officially proven yet
So what should builders and operators watch next You should watch how AI companies fundamentally adjust their physical security protocols moving forward Watch to see if companies start aggressively stripping all partner branding and logos off their physical assets to avoid drawing unnecessary attention The logistics of basic research and development just got a lot more complicated You can't just park a branded truck in a lot anymore It is a gritty frustrating reminder that the AI boom isn't just happening in the cloud It exists in the messy physical world where people will steal a truck just to see what's inside So true Meanwhile in the realm of the
hardware actually running these models OpenAI has officially opened faster access to GPT 6 Astra The new speed tier runs on NVIDIA's Blackwell GPUs And they are claiming massive token generation speedups This is a major development for the developer community Let's lay out the specifics of this rollout OpenAI is officially calling it GPT 6 Astra Ultrafast It is now available right now via their developer API And it is also available to eligible enterprise users on chat GPT work and the Codex programming assistant The headline metric driving the excitement comes directly from NVIDIA They claim this new Ultrafast mode generates tokens up to eight times faster than
the standard Astra deployment Eight times faster is not an incremental update That is a massive generational leap But my background tells me we need to look closer at how they actually achieve this It is never just about plugging in faster hardware You are exactly right You don't get an 8x speed up just by racking new servers The crucial detail here is deep co optimization Explain how that optimization actually works Because OpenAI didn't just buy the new Blackwell chips and run their old standard inference software on them No they didn't According to NVIDIA's technical blog OpenAI actually utilized its own internal proprietary AI models to test
and implement massive improvements to the inference software serving Astra Oh wow So they use their own AI to write the software that optimizes the AI Yes They specifically targeted and improved the high performance kernels Okay we need to break that down What exactly is a GPU kernel in this context Think of a kernel as a highly specialized microprogram It is the specific set of instructions that tells the graphics processing unit exactly how to perform the complex matrix math required for neural networks It dictates how the chip crunches the numbers Exactly When a model generates text it is essentially moving massive amounts of data back and
forth between the GPU's memory and its processing cores This is often a memory bandwidth problem not just a pure calculation problem Let me use an analogy here Think about a restaurant kitchen The GPU processing cores are your fast highly trained chefs Okay The GPU memory is the pantry where all the ingredients are kept That is a perfect way to visualize it If the kitchen has a terrible layout your fast chefs spend most of their time walking back and forth to the pantry to grab one onion at a time Right Chefs are fast but the food takes forever to cook The system is bottlenecked by the
physical layout The kernels are the kitchen layout Right OpenAI use their internal models to redesign the kitchen They rewrote the kernels so the data flows perfectly from the memory to the cores without any wasted movement The Blackwell chips are incredibly fast chefs but the co optimized software is what allows them to actually cook eight times faster It is a relentless ongoing process of tuning the entire inference stack Hardware and software have to evolve together So why does this specific 8x speed up matter so much right now Why are developers so focused on raw token generation speed Because the industry is shifting fundamentally toward agentic workflows
Let's talk about agents A standard chatbot is basically a fast reader You ask it a question it reads its training data and it answers you A single interaction A single output But an AI agent operates in a continuous loop Think about an AI coding assistant trying to build a website It doesn't just write one answer Right It writes a block of code Then it has to read the result of that code Then it has to run a test Then it has to call an external tool like a compiler Then it has to check the output of that compiler Then it writes a fix It is
a massive loop of constant back and forth interaction Plan execute evaluate refine And the critical bottleneck in that loop is the waiting time Every time the agent pauses to generate the next step in its thought process the whole system stalls If an agent takes 10 seconds to generate its next plan of action and it has to plan 50 times to complete a task you are waiting a long time But if you can generate tokens 8x faster you dramatically shrink those edit test and debug cycles The waiting time basically evaporates A complex multi step task that used to take five minutes of agonizing pauses might now
execute in 40 seconds That speed fundamentally changes the viability of real time interactive applications It turns a clunky novelty into a fluid tool But we need to bring in the dry analytical caveat here because 8x faster token generation does not mean the agent finishes its job 8x faster That is the reality check developers need to understand The 8x metric is strictly for token generation It is a raw measurement of how fast the model spits out text or code strings once it starts processing It is the speed of the chef chopping the onion Yes but it is not a guarantee that a complex multi step task
will finish exactly eight times sooner overall Because there is still massive latency everywhere else in the system If your coding agent has to make a network call to an external API to pull weather data it still has to wait for that external server to respond Exactly Network latency doesn't care how fast your GPU is API limits database queries external tool execution times those remain exactly the same The model generates text faster but the agent still has to wait on the real world And we also need to note that NVIDIA's claim is an up to maximum metric It represents peak performance under ideal conditions It is
not a sustained minimum speed you will hit on every single prompt There is also a secondary narrative here involving Amazon Amazon Web Services announced their own access to Astra Ultrafast on the Amazon Bedrock platform but they are using entirely different performance numbers in their marketing Yes AWS attributes a different claim to OpenAI In their press release AWS cites up to six times faster API inference and up to 300 tokens per second So developers are looking at an 8x claim from NVIDIA and a 6x claim from AWS for the exact same model update Those figures are simply not interchangeable They are measuring different aspects of the
data pipeline under entirely different hardware networking conditions Okay AWS is measuring the end to end inference speed over their specific cloud infrastructure NVIDIA is likely measuring raw silicon performance on a dedicated cluster You should not collapse those numbers into one universal speed guarantee Your mileage will absolutely vary based on where you are actually running the model So what should the developer community be watching next Watch how the actual access rolls out Right now the chat GPT work and codecs availability is limited to specifically eligible users It is not wide open to the general public yet More importantly watch the developer forums We need to see
whether these theoretical token speed ups actually translate to noticeable workflow efficiency for engineers building products in the wild Right The theoretical speed limit is undeniably higher but we need to see the practical impact on daily productivity Exactly Shifting from hardware speeds to model training researchers have introduced a new method called Pivot OPD The approach targets a very specific problem in multi turn AI tasks making an early mistake and completely failing to recover from it This research is a fascinating look under the hood of how AI agents actually behave when they get confused The paper was just submitted to RxVIF on September 30th The research team
ran extensive tests on various sizes of the Quen3 model family They specifically looked at complex tasks that required the AI to take multiple sequential steps to reach a goal Like navigating a file system or debugging a long script Exactly And the researchers found something stark in the data In over half of the failed runs the AI made what they termed a pivotal mistake very early in the process A pivotal mistake is an action that fundamentally moves the agent farther away from completing its goal rather than closer to it And once the agent makes that first bad move the entire situation degrades rapidly The agent is
now operating in an altered environment Its context window is essentially poisoned by its own error Right Its subsequent decisions are all based on that first incorrect assumption It starts hallucinating fixes for a problem it created It spirals and it dooms the rest of the task completely But the researchers found a very optimistic data point buried in those failures These wrong turns were often highly recoverable The situation wasn't fatal No If a human stepped in and guided the model for just one or two turns immediately after the mistake they could easily put the agent back on track to success Yeah The model had the capability to
finish the job It just didn't know how to back out of a dead end That is exactly where the new PivotOPD method comes in The training approach pairs a smaller student model with a highly capable teacher model And the training focuses on two separate vital lessons prevention and recovery Let's walk through the prevention phase first The student model is generating actions in a simulated environment When it makes that early pivotal mistake the teacher model pauses the simulation The teacher steps in and provides the exact gold action It shows the student what it should have done at that critical junction It corrects the specific error But
it does not stop there The second lesson is far more important The recovery phase The teacher allows the simulation to proceed from the mistake Then the teacher supplies the correct recovery actions for the next few turns Based entirely on the messy broken situation the student just created Right It explicitly teaches the AI how to dig itself out of a hole of its own making It's brilliant Now technically the researchers achieve this using two mathematical concepts They use reverse KL divergence to steer the student away from the initial error And they use forward KL divergence to teach the recovery behaviors We need to break those concepts
down Sure KL divergence is essentially a way to measure how one probability distribution differs from a second reference probability distribution OK stop right there That is absolute textbook definition jargon We need an explain like I'm five analogy here Let's use the bowling alley I like the bowling alley Let's do it Think of reverse KL divergence like putting up the bumpers in the gutters of a bowling lane OK When the AI is about to make a mistake reverse KL acts aggressively It forces the model down a very narrow specific path The exact gold action the teacher provided Exactly It prevents the ball from going in the
gutter by physically blocking off the bad options It demands narrow conformity It is highly effective for mode seeking behavior which means finding the single most correct answer And what about forward KL divergence for the recovery phase Forward KL is entirely different It is more like giving the AI a wide angle lens When the AI is stuck in a bad situation there isn't just one single way out Right There might be five acceptable ways to recover Forward KL encourages the student model to look at the broad distribution of all the viable recovery behaviors the teacher suggests It teaches flexibility Exactly It transfers a broad range of
acceptable actions letting the student learn how to adapt dynamically rather than just memorizing one strict path The results of combining these two approaches are actually pretty compelling The researchers ran tests across three different established agent benchmarks Pivot OPD showed the strongest average performance against 13 different baseline training methods They saw a concrete 5 5 percent improvement on the ALF world benchmark using a relatively small QEN 3 1 7 billion parameter student model And they made sure it wasn't just a quirk of the QEN architecture They also tested a completely different model family They used a Nemotron 3 5 student model on the SWE Bench verified
benchmark SWE Bench is a rigorous software engineering test It requires the agent to resolve real GitHub issues And they saw a solid 3 2 percent increase in total tasks resolved using the Pivot OPD training That proves the training method transfers across different architectures and entirely different task domains I hear the enthusiasm but I have to push back a little on the scale of these numbers A three to five percent bump in performance sounds nice on an academic paper but is this just a hyper specific optimization trick to win leaderboard points That is a fair skepticism Let's think about a car's GPS navigation system Okay When
you miss your exit on the highway the GPS doesn't just crash It says recalculating Right It figures out how to guide you to the next exit have you do a U turn and get you back on the route It doesn't just tell you to keep driving straight into a lake That is a perfect analogy for this research The baseline AI models are currently driving straight into the lake They make a wrong turn and they just keep accelerating in the wrong direction Because they do not know what to do after a missed exit Pivot OPD is teaching the AI the concept of recalculating It gives them
the mechanism to pause assess the error and execute a U turn But the primary limitation is right there in the benchmark data These performance gains are currently tied to very specific controlled models and highly structured benchmarks Right Off world is essentially a text based game environment with strict rules Sui Bench is highly structured coding logic The open question for you to watch is whether this prevention and recovery approach actually scales broadly Will it improve general agent reliability in unstructured chaotic real world workflows We will have to watch if the major frontier labs adopt this specific dual training approach for their next generation models Teaching an
AI to apologize and fix its own code might be the missing link for true autonomy All right we are going to move into our quick read section now We will keep these brief but distinct In our first quick read a man in Singapore is facing serious criminal charges over a single AI generated image Wait really Over one image Yes Yi Lin a 30 year old man used Google's Gemini platform to create a highly realistic fake image of a large crocodile The image was specifically prompted to look like it was taken at the local Pandan Reservoir On August 20th he shared this fake image in a
private message with a single colleague But the image didn't stay private No It spread rapidly through local chat groups It caused a massive public panic The National Water Agency took the threat seriously and actually suspended all rowing and canoeing activities at the reservoir while they searched the waters for this non existent crocodile Once the authorities determined the image was a total fabrication they tracked it back to him and reported him to the police He is now facing two major criminal charges The first is for transmitting a false message that caused public alarm The second and much more severe charge is for obstructing justice The police
allege that when he realized he was under investigation he deliberately deleted the fake image from his phone Furthermore he allegedly deleted the entire Gemini app from his device in an attempt to conceal his actions from the police The potential penalties are severe The obstruction charge alone carries a maximum penalty of up to seven years in prison The false message charge carries up to three years He reportedly plans to plead guilty in court this November This case aggressively highlights the massive real world public costs and the direct criminal liability of fabricated imagery Even if you only send the picture to one single person you are responsible
for the panic it causes In other news an interesting test of AI image boundaries was conducted by reporters at The Guardian They specifically asked several major AI chatbots to edit a synthetic image of a woman wearing a hijab The tests revealed wildly different safety boundaries and internal logic across the major models The base image was entirely AI generated so it was not a real identifiable person whose privacy could be directly violated Right When asked to alter the image ChatGPT and Elon Musk's Grok both completed the direct request without hesitation They simply generated a new version of the image removing the hijab entirely Google's Gemini responded
with a much more complex set of rules It initially refused the direct request to remove the garment citing its internal consent and safety rules regarding people But when the reporters changed the text prompt to simply ask the model to make the synthetic woman look more western Gemini complied immediately It removed the hijab It was the exact same requested visual change It was just achieved through a phrasing workaround Anthropic's Claude took a different path entirely It stated clearly that it lacked direct image editing functionality in its current build but it added a philosophical boundary Claude stated that even if it possessed the technical ability to edit
the image it would flatly refuse the request The model explained that altering religious dress could embarrass or misrepresent a person's core identity OpenAI later acknowledged to the reporters that exposing the hair of someone who wears a hijab in real life could violate privacy norms or religious practices The test sparked a massive industry debate on how automated AI image editors should navigate the distinct sensitive boundaries between basic pixel manipulation and deep religious identity Next up some significant personnel news out of OpenAI The company has officially parted ways with three researchers from their safety team following a quiet internal investigation The details surrounding these departures are incredibly
murky A detailed report from the Wall Street Journal alleges that the three researchers shared highly confidential company information with an outside AI safety organization But OpenAI's official public statement is much narrower in its phrasing A company spokesperson simply stated that the individuals mishandled sensitive information outside of established company procedures They cited policy violations and a fundamental breach of trust as the reason for termination What remains entirely unknown to the public is what specific material was actually shared Was it technical waddle weights Was it internal safety debate memos We simply do not know We also do not know if these researchers utilized internal reporting channels first
before allegedly taking the information outside the company The entire narrative of whether this was heroic whistleblowing or corporate espionage hinges entirely on those missing facts which neither side has released We are going to move to three takeaways from today's conversation First across the board from virtual clothing try ons and chat GPT to the acceleration of multi turn agent workflows on Blackwell chips The AI industry is aggressively moving past the basic discovery phase They are no longer just answering single questions They are actively building the infrastructure for sustained continuous engagement loops Second we are seeing the acute growing pains of both physical and digital security We
have autonomous tech startups dealing with 50 foot trailers full of sand being targeted by thieves in the real world At the same time frontier labs are navigating complex confidential information boundaries with their own internal safety researchers The security perimeter is failing in both domains Third context is everything for AI models right now Whether we are talking about prompt engineering loopholes in image editors that reveal internal biases or training agents using KL divergence to recognize their context has changed so they can recalculate after a bad first step the models have to understand where they are So what is one development we should be watching tomorrow Watch
the continued push by researchers to force models to correct their own reasoning mid task The pivot OPD research we discussed is just the beginning We are going to see a flood of papers about models learning to step back evaluate their own context window and fix their own errors in real time without human intervention And we want to leave you with a final thought to mull over If we can successfully train an AI agent to recognize its own pivotal mistakes and course correct on the fly what happens when it starts applying that exact same real time critical scrutiny to the prompts you feed it It is
a fascinating shift in the power dynamics of how we interact with software For more deep dives and daily signals head over to superpowerdaily com Thank you for listening We'll see you tomorrow
Original reporting
Read the complete Superpower Daily coverage behind this episode, including reporting context and source links.