Loading page…
Loading page…
The Signal / Superpower Daily
AI’s rewards and access are uneven: women remain underrepresented in new hires, while Google is narrowing Gemini’s lower tiers. Today’s lineup also brings a qualified accounting win for AI, deeper Claude Code customization, and contrasting fortunes for chip stocks.
Superpower Daily: The Signal
Episode guide
AI’s rewards and access are uneven: women remain underrepresented in new hires, while Google is narrowing Gemini’s lower tiers. Today’s lineup also brings a qualified accounting win for AI, deeper Claude Code customization, and contrasting fortunes for chip stocks.
Full transcript
Select any transcript timestamp to continue listening from that point.
Welcome to The Signal from Superpower Daily with Maya and Theo We are diving into a massive 45 000 hidden gap in the artificial intelligence boom today Right Well you're looking at why the human infrastructure behind these frontier models is really struggling to scale So let us start with the human reality behind all this rapid scale The AI sector is obviously generating wealth at an unprecedented rate right now Yeah absolutely unprecedented But you know the distribution of that opportunity is completely fractured New data reported by The Guardian citing research from LinkedIn reveals a really massive hiring disparity Women accounted for only about 25 of new AI
hires over the past year And that number requires immediate context If you look at the broader economy outside of this specific sector women made up 50 of new hires during that exact same period Wow Okay so that is double Exactly We're watching the industry that is actively building our next generation of technology operate with a severe structural imbalance from day one Well the strategic implications here are enormous LinkedIn found that AI job postings have essentially doubled since 2023 And the compensation premium is staggering Roles involving these technologies pay more than twice as much on average as roles without them But you have to look at
where people actually sit within the organizational chart That pay premium is definitely not shared evenly across the workforce Meaning what exactly Well women in this sector are disproportionately concentrated in lower paying roles They are heavily represented in support project management and data annotation Right the operational stuff Exactly They are vastly underrepresented in core machine learning engineering and architecture That occupational divide just completely skews the economic reality of the boom I mean across all AI occupations the median pay for men is a staggering 45 000 higher than the median pay for women We need to be very precise about what that specific data point means though
Okay go ahead That 45 000 gap is a reflection of the types of jobs people hold It is not a direct measure of unequal pay for identical technical work Men are landing the highest paying technical roles Women are landing the lower paying operational roles Ah I see We also have to look at the risk profile outside of the tech industry itself The data shows women are more likely than men to hold jobs that are highly exposed to automation disruption Customer service and administrative functions are prime examples there And that introduces a crucial caveat to the data High exposure to disruption is a measure of vulnerability
to changing work It means the job is going to change right Right It is a forecast of shifting responsibilities It is not a definitive tally of jobs that have already been eliminated A highly vulnerable job might evolve significantly rather than disappear entirely But the burden of that transition is clearly skewed toward women So we really need to unpack why this hiring gap exists right now Brenda Darden Wilkerson is the president of AnitaB org and she points directly to the speed of hiring as a primary culprit here The current environment is a pure land grab Companies are racing to scale up as fast as mathematically possible
Yeah everybody is just running And when you compress a hiring cycle you naturally default to the path of least resistance You rely heavily on familiar networks You lean on employee referrals That creates a compounding network effect does it not I mean if your founding engineering team is mostly male their immediate professional networks are often mostly male Exactly Fast hiring bypasses deliberate recruitment It completely circumvents any effort to build a diverse pipeline We saw this exact same pattern during the Web 2 0 boom Are we just baking the exact same structural biases into the foundation of a brand new technological era because we refuse to slow
down I mean that is the classic pipeline problem But the pipeline is only half the story here We have to address the retention problem Retention is a huge issue Right Getting hired into a frontier lab is only the first hurdle Surviving the culture is a completely different challenge Jayita Puttatunda is an engineering lead at the investment firm Turing Her experience perfectly illustrates this retention crisis She noted that being the only female engineer on a technical team is incredibly common And she points out that mentorship is practically non existent in these hyper growth environments Nobody has time for it Exactly The culture fundamentally demands 12 hour
workdays It requires weekend sprints to hit arbitrary release schedules That kind of relentless unstructured pressure actively filters out diverse talent post hire It heavily favors people with zero outside responsibilities Right And Puttatunda shared a very specific experience about returning from a four month maternity leave Four months in this industry is an absolute eternity It really is She returned to work and found that the entire technical stack had shifted Oh wow Everything Everything had changed The frameworks were completely different The dominant models had changed She described the process of catching up as completely overwhelming I cannot even imagine Well that is the brutal reality of the
current technology cycle The pace of change is punishing If you step away for a single financial quarter your architectural knowledge might be entirely obsolete She managed to catch up though But she explicitly credited supportive colleagues and a husband who shared child care responsibilities That is key That support system was the only thing that made her return viable Her story proves a critical point I think Women are not leaving these technical roles because of a lack of ability Right They are leaving because the industry refuses to build support systems that acknowledge the impossible pace of its own development And you know we see these exact same
structural barriers on the entrepreneurial side Urvashi Batra is the co founder and CEO of a platform called PriorityWise She stated on the record that investors simply take her less seriously than her male co founder Yes She said they are demonstrably more likely to secure venture capital investment when he is the one delivering the pitch That is an incredibly frustrating reality for female founders trying to build in this space It is deeply frustrating But we do need to include an important caveat regarding the founder experience here What is that These are subjective accounts of individual fundraising efforts They highlight very real cultural pain points However they
are anecdotal experiences They are not measured statistical explanations for the broader engineering hiring gap we discussed earlier That is a fair point The structural barriers are real and they are embedded in the daily operations of these companies though Absolutely So what should you watch next Watch the hypergrowth companies closely Watch to see if they actually evolve their hiring practices past those closed referral systems You should also watch how the industry adapts to support retention Will companies implement paid catch up time for engineers returning from extended leave The industry must eventually build operational frameworks that allow humans to actually work in this space sustainably Moving on
to AI in the workplace Startup Mercor found its AI models can now beat licensed human accountants on simplified month end tasks This research was published on October 1st They tested Frontier models on four specific month end accounting scenarios and the models completed the assignments perfectly They achieved a 100 score And they tested these models against 12 licensed certified public accountants These were not entry level bookkeepers fresh out of college either No definitely not These CPAs had an average of five and a half years of professional experience And the models beat the human accountants in speed They beat them in accuracy And the cost efficiency was
just staggering How staggering The computing cost per grading criterion met was more than 10 times lower than the human labor cost Okay I need to push back on this a little bit We are talking about CPAs with over five years of experience losing to a model What exactly were they tested on Anyone who has ever run a month end close knows it is usually totally chaotic That is exactly the right question to ask Each of these four assignments was highly structured The tasks required searching a provided set of company working files They required finding specific figures within those files They required performing standard calculations And
finally they required delivering a cleanly formatted table of results Okay so it is a closed loop search and generation task Right But this matters significantly because it proves an immense capability in data heavy assignments The models are mastering the mechanics of file retrieval and calculation workflows And the historical context here is what really matters Just 18 months earlier the best available models fell completely flat on these exact same tests Really 18 months ago Yeah Wow The human average score on these tests was about 37 back then The leap in performance on self contained structured work over a year and a half is staggering The test
authors actually built realistic traps into these assignments too The requirements were hidden deep in the files The calculations compounded on one another Right A single missed number early in the process would just ruin the final deliverable They expected mid level accountants to score around 55 And the models navigated every single one of those traps flawlessly But this brings us to the most important limitation of the entire study Which is The test conditions were heavily stripped down to create a sterile environment Ah so the human accountants were operating in a total vacuum Exactly They had no co workers to ask for help They had no client
conversations to clarify ambiguous data They had no accumulated institutional knowledge about the company they were supposedly working for Well Mirko points out that these sterile conditions actually put the human accountants at a severe disadvantage Real accounting is not just math No not at all It relies heavily on institutional context It relies on knowing which specific co worker has the context for an unexplained expense Right And that sterile environment is exactly why these models are not just doing everyone's corporate taxes right now Yeah corporate accounting is incredibly messy Half the job is chasing down a stubborn fender for a missing receipt The model in this test
never had to negotiate with a human being Exactly And there is another major caveat to how this study was designed Okay The researchers originally set out to measure how much these models could assist a human accountant They wanted to test a combined workflow But the models achieved perfect scores completely on their own They hit the ceiling of the benchmark Right Because they scored perfectly there was no room left to measure human improvement The study was forced to compare unassisted humans against models working completely alone Wow We also have to zoom out and look at the broader picture here These four simplified tasks are just a
tiny fraction of a much larger test called the Apex Accounting Benchmark And that full benchmark contains 160 tasks spread across 10 simulated companies The full benchmark is a massive jump in complexity It includes reconciling difficult accounts It includes accruing complex expenses It involves posting live transactions and producing comprehensive reports And the models absolutely do not score perfectly on the full test Not even close right No The Dakota reported on the broader benchmark results Claude Opus 5 5 led the entire pack but it only met 61 8 of the total grading criteria That means almost 60 of the tasks in the full benchmark have not been
fully solved by any model on the market Right That is a massive reality check Hitting some grading criteria on a simplified test is entirely different from autonomously completing a messy real world corporate assignment What should you watch next Watch to see if these models can handle wider unstructured responsibilities They need to prove they can manage the messy human context required to actually run an autonomous accounting close They have to move beyond passing isolated benchmark tests in a vacuum Next up a different kind of tech gold rush Slovenia's si domain registry is seeing an unprecedented surge in registrations amid Donald Trump's AI rebranding effort This is
a fascinating look at the speed of internet speculation The country code web address suffix for Slovenia is si Right The organization that manages these domains is called Register si and they reported a massive unexpected spike in activity throughout September Clara Hermann is a spokeswoman for the registry She explicitly called the pace of these new registrations unprecedented September saw up to 46 066 registrations On September 30th alone they recorded 11 000 new registrations in a single day To put that volume in perspective an average month like August saw just a few thousand registrations total That is a huge jump It really is And this massive surge
coincides perfectly with a specific political directive in the United States U S President Donald Trump is pushing an initiative to replace the term artificial intelligence in federal communications His preferred replacement term is super intelligence The acronym for that is si Right And this matters far beyond the political headline because it highlights the hyper speculative nature of the tech market Buyers are rushing to secure a foreign country code suffix They are placing a massive bet that this new label catches on globally Exactly If the broader tech industry actually adopts the si acronym those domain names instantly become highly valuable digital real estate It is wild to
watch internet speculators jump on a single government memo and buy up tens of thousands of Slovenian domains in a matter of hours It is But we need to kind of clearly explain the mechanics of this behavior This is pure domain investing It is not automatically cyber squatting OK clarify that What is the difference Cyber squatting is a specific legal term that implies acting in bad faith It usually involves buying a trademark name specifically to extort the rightful owner Like buying a big brand's name before they can Right This is a completely different strategy This is betting on a broader language shift It is exactly like
buying cheap land in a rural town because you heard a rumor that the state might build a new highway through it You are just hoping the traffic eventually comes your way Exactly The risk profile is entirely unique Speculators are betting that a localized government directive will fundamentally alter global tech branding But there are several important limitations and caveats to this story First we need to clarify the exact scale of the surge There are two different reporting tallies floating around The BBC reported 44 000 registrations for the month TechRadar reported over 46 000 Right And TechRadar ran a headline claiming an over 100 times increase But
that math compares one exceptionally busy single day to an average single day in August So it is not a 100 times jump for the entire month No definitely not Both tallies absolutely confirm a massive jump But the hyperbole in some of the reporting requires a reality check The most important caveat involves the actual scope of the political directive itself The executive directive only applies to U S executive branch agencies It dictates the vocabulary for official government communications and federal contracts It absolutely does not require private businesses to adopt the acronym It does not force the wider tech industry to change its terminology A federal instruction
for government paperwork is not a mandatory industry rename Exactly There were also wild rumors circulating that this surge would bring millions of windfall revenue to the Slovenian government Oh right The registry stepped in and quickly denied that They charged registrars a flat fee of just 10 euros per domain They stated clearly that those windfall claims have zero basis in reality What should you watch next Watch whether private enterprise actually adopts the acronym in their own marketing and product launches Look beyond government documents If the wider tech industry simply ignores the rebrand and sticks with the current terminology these speculative domains may end up completely worthless
Meanwhile over at Anthropic the company just gave developers a radical new way to customize Cloud Code by adding feature replacing mods This represents a major structural shift in how coding agents operate locally On October 1st Anthropic introduced mods Okay These are small custom TypeScript functions They attach directly to specific events inside the Cloud Code execution loop These mods are incredibly powerful They can completely rewrite agent behavior on the fly They can replace built in interface features They can intervene in tool calls and override permission requests before they ever execute And this goes far beyond the existing customization options developers are used to Previously developers relied
on standard hooks Hooks just that they're in wait right Exactly Hooks are entirely reactive They can listen for an event and respond after it happens But they cannot fundamentally change the event itself Mods change the paradigm completely They are deep active interventions A mod can run before an event triggers It can run after an event finishes It can completely replace the event and run instead of it This matters immensely because it shifts ultimate control from the model provider back to the local developer You could write a mod to block a specific tool call Right You could force a failed tool call to retry with different
parameters You can intercept a tool's output and strip sensitive secrets from the data before Cloud ever sees it You can even build custom user interface buttons directly into the terminal Anthropic has already migrated their own built in diff feature over to this new system Right Users can now disable the official diff feature entirely and install their own custom version if they prefer a different visual output This points toward a very specific future architecture for AI agents does it not The core models will become incredibly small fast and lightweight Organizations will simply custom build their exact proprietary agent workflows on top of that core using extensive
libraries of mods Wait this sounds exactly like what happened with video games like Skyrim Bethesda built the base game but the community practically rebuilt the entire interface added new mechanics and fixed bugs using mods That is a perfect analogy The community ended up building better tools than the original developers Are we saying Cloud is just becoming the base engine for enterprise AI That analogy holds up perfectly It creates a massive vibrant ecosystem of community driven tools But video game mods also bring massive systemic risks Oh right Security That leads us directly to the most important caveat of this new system The security trade off is
absolutely enormous These mods do not run in an isolated sandbox When you install a mod it inherits Cloud Code's full access permissions to your local machine That is a staggering level of access to hand over to a small TypeScript function It really is If you install a malicious or poorly written mod it literally has the keys to your entire system A lack of sandboxing is the ultimate double edged sword It provides incredible developer speed and limitless flexibility but it completely sacrifices enterprise security Anthropic is advising users to only install mods from trusted sources but anyone who has ever worked in open source knows exactly how
that goes Oh absolutely Developers will blindly copy and paste code from a random repository to solve a quick problem without ever auditing it for vulnerabilities And load order also introduces a massive vulnerability here When several different mods attach to a single event the first one loaded sees the event first It has the power to alter the data before any other mod sees it It also receives the final result last giving it the final say on the output That is wild Well Anthropic is trying to mitigate this massive risk for larger organizations For their team and enterprise plans they forcibly enforce a protective mod called SecDefault
Right This built in security mod is designed to load first It blocks highly risky actions like a user trying to override a system permission denial Network administrators can put their own custom security mods ahead of it in the load order but they are strictly required to keep SecDefault enabled in their active list What should you watch next Watch exactly how organizations handle the security governance of these mods Installing unverified mods could quickly become a massive vulnerability for enterprise networks The tension between demanding extensive developer customization and maintaining strict corporate security will completely define this next phase of AI deployment We are shifting now to our
Quick Reads section First up Google is cutting Gemini model access for its lower subscription tiers This change starts on October 9th Free tier users will lose access to the standard Flash model They are being relegated to the smaller Flash Lite model Google AI Plus subscribers will also take a significant hit They will lose access to the more capable Gemini Pro model later in October Google stated they will email Plus users with their specific cutoff date soon This matters because it forces a hard decision on the user base Users who rely heavily on the Pro model for complex multi step tasks will be forced to upgrade
They must move to the significantly more expensive Google AI Pro or Ultra subscription plans The key limitation here is understanding what is actually changing Google AI Pro and Ultra plans are completely unaffected by this move This is a strategic restriction on which subscription plans get access to which models It is not a withdrawal of Gemini Pro from the service entirely Plus users will still keep their expanded storage benefits They will also keep access to the Flash Lite model The looming question for Google is whether those remaining benefits justify the monthly subscription cost once the Pro model is removed Next up startup Tavis has previewed a
new video call AI called Gryphon Lite In a one minute company test 26 out of 54 people completely mistook the AI for a real human being That is nearly half the participants failing to identify the model Tavis notes this represents a massive jump in capability from their previous system In tests of the older version only one out of 41 people mistook it for a human This matters because Gryphon Lite is being built as a human interaction model It does not just generate a static realistic face Right it is active It listens speaks and gestures simultaneously It processes your vocal tone and facial expressions live and
reacts to them in real time The major caveat here is the specific design of the test This was an incredibly small sample size Very small More importantly users went into the test expecting to speak to a human Furthermore the interaction was strictly limited to one minute A brief one minute impression is entirely different from sustaining a coherent conversation over a 10 minute meeting We also have to look at the benchmark data from NVIDIA Gryphon Lite scored 3 73 on perceiving conversational cues That score still lags significantly behind the standard human reference score of 4 20 And there is no public launch date set yet The
model remains strictly restricted to trusted testers while the company works on implementing necessary safety and disclosure features Finally Anthropic is making a massive investment in human capital They are spending 100 million to launch the Claude Frontier Academy The stated goal of this academy is to train 10 000 enterprise engineers They want these engineers fully certified and ready to deploy Claude inside massive businesses by the end of the year 2027 This matters because it addresses a critical unglamorous bottleneck in the industry Taking a model from a cool chat interface prototype and turning it into a secure working business system that handles sensitive data is incredibly difficult
Right And the academy curriculum involves rigorous simulated enterprise deployments Participants must pass practical hands on assessments And then what They then enter a 12 week residency where they lead a real deployment project at their actual employer The limitation to watch here is the timeline and the access model The 10 000 figure is a lofty goal for the end of 2027 It is not a current count of available graduates ready to work today The first final credentials from this program will not be issued until early 2027 Oh wow Additionally this is not an open enrollment boot camp for junior developers It is strictly by corporate nomination
aimed entirely at experienced engineers already working inside partner organizations We are moving now to our three takeaways from today First the human capital gap is widening severely The industry is creating massive wealth and driving incredible efficiency Job postings are doubling and the pay premium is extraordinary But structural barriers fast paced referral hiring and punishing retention cultures are actively keeping women out of the most lucrative technical roles Second agent autonomy is rapidly expanding beyond basic chat We saw Mercor's models achieve perfect scores on isolated month end accounting tasks We saw Anthropic introduce deep machine access mods for local code agents The technology is aggressively moving from
a reactive assistant to an integrated autonomous digital worker capable of executing complex workflows Third the reality check of enterprise deployment remains incredibly severe We see rapid wild speculation like buyers hoarding Slovenian domain names based on a single government acronym But the actual work of implementation is grueling Anthropic has to spend 100 million just to teach experienced engineers how to safely integrate AI into real corporate environments The gap between speculative hype and practical business integration is still massive Watch how the open source community responds to Anthropic's new mod system tomorrow Specifically watch if independent developers build a community sandbox to fix those glaring security gaps before
enterprises refuse to adopt it You can find all this at superpowerdaily com Thank you for listening We'll see you tomorrow
Original reporting
Read the complete Superpower Daily coverage behind this episode, including reporting context and source links.