Latent Space in 3 minutes

Unofficial daily recap of Latent Space. Each recap links to the original episode so you can listen to the ones that grab you! More recaps at https://thedaily.fm

Listen to the original podcast →

Cadence: On demand
Length: 3 minutes

Subscribe, Combine, Customize

Subscribe to this podcast
?Receive all episodes to this podcast in the apps below or anywhere that supports RSS.
Combine these episodes into your pod
?All episodes from this podcast will be fed into your own.
Sign up to add to your own podcast
Customize this pod with your own sources
?Use this if you want a brand new podcast with its own episodes using different sources.
Sign up to customize this pod

Sources

  • pod:api.substack.com

Episodes

Latent Space in 3 minutes: 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
Created: August 11th, 2026 - 14:15 PT
Script

Here is The Daily FM summary of the Latent Space that aired on Tuesday August 11th. This episode explored what Chai Discovery calls a “phase shift” in BioAI: moving drug discovery away from slow, brute-force experimentation and toward precision engineering with AI models.

Chai cofounder Matt McPartlon and platform leader Neil Patil described the company as a software and modeling layer for pharma rather than a company trying to build its own drug pipeline. Their pitch is that better protein-structure and protein-design models can help drugmakers find therapeutic candidates faster, with more control over where and how they bind. Chai has partnered with Eli Lilly, Pfizer, Novartis, and argenx, and argues that its success is tied directly to whether those partners make better medicines. [1]

The central technical focus was antibodies: Y-shaped immune proteins whose tips can be engineered to attach to disease targets. Traditionally, antibody discovery often means immunizing mice or screening billions of possibilities in yeast-display experiments, then hoping a few molecules stick. Chai’s models aim to generate candidates intentionally, including molecules that bind a specific site, avoid similar proteins that could cause side effects, or bind both human and animal versions of a target to support preclinical testing.

McPartlon described Chai 1 as a structure-prediction model: given a protein sequence, it predicts the 3D shape. Chai 2 made the bigger leap into design, generating both a candidate’s sequence and structure for a target. In a high-profile internal challenge, the team designed antibodies against 50 targets and got binders for roughly half. One remarkable validation result showed a predicted structure only 0.33 angstroms from cryo-EM measurement—about one-third the width of an atom. The team initially suspected the lab had accidentally sent their own prediction back.

But the guests stressed that binding is only the beginning. Good drugs also need strong affinity, selectivity, stability, safety, manufacturability, and the ability to avoid unwanted aggregation. Chai 3 and subsequent iterations focus on pushing toward therapeutic-grade molecules rather than treating hit discovery and lead optimization as separate, years-long waterfall stages. [2]

Patil argued that the product should not look like a chatbot. Instead, Chai built something closer to Figma, SolidWorks, or Photoshop for molecules: scientists can visually select an epitope—the precise target region—and ask the model to generate binders under constraints. The product must also meet pharma’s strict security and intellectual-property requirements, including isolated customer deployments. [3]

A moving moment came when Patil recalled a pharma scientist crying after Chai helped generate a binder for a target she had worked on unsuccessfully for ten years. That illustrated the broader claim: these systems are beginning to work in real campaigns, not just on benchmarks.

The biggest remaining bottleneck, McPartlon said, is rapid experimental validation. Models can generate hypotheses quickly, but wet-lab proof still takes weeks or months. Patil’s answer was talent: BioAI needs more engineers and researchers to realize that biology is becoming computationally accessible. Their final message was ambitious but clear: as models predict structures within atomic accuracy and reliably design binders, biology may become less like feeling around in the dark and more like an engineering discipline. Thank you for listening to Latent Space in 3 minutes from The Daily FM. See you next time! [4]

Source Evidence
  1. Latent Space: 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
    ...nd optimization where each of these has a gate and takes a few months to a few years is this very like waterfall mod where the cost of trying things and getting things early is very expensive. But I think to what Matt's saying, if you start to get in a regime where you can have models give you really promising candidates, you can start to make that look a lot more like a loop. Right. It's akin to becoming more agile in software development. But now the next problem is agonists. How do you reliably one shot hitting a switch on a cell. Right. Or buy specifics or ADCs. Right. And I think this levels of abstraction that we're going to have to climb with the product as like the models get better. If you have like these really good primitives for structure prediction and binding and design and you can kind of compose them then you can start to just like grow into like the o...
  2. Latent Space: 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
    ...of load up your molecule, there's this almost like Photoshop esque like design suite. You have this equivalent of paint tool to kind of paint your epitope. You have this equivalent of a content aware fill tool to kind of get your, your binders generated from Chai. And I think to add to that, right? Yeah. This notion of target discovery and hit discovery and optimization where each of these has a gate and takes a few months to a few years is this very like waterfall mod where the cost of trying things and getting things early is very expensive. But I think to what Matt's saying, if you start to get in a regime where you can have models give you really promising candidates, you can start to make that look a lot more like a loop. Right. It's akin to becoming more agile in software development. But now the next problem is agonists. How do you reliably one shot hitting a...
  3. Latent Space: 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
    ...of load up your molecule, there's this almost like Photoshop esque like design suite. You have this equivalent of paint tool to kind of paint your epitope. You have this equivalent of a content aware fill tool to kind of get your, your binders generated from Chai. And I think to add to that, right? Yeah. This notion of target discovery and hit discovery and optimization where each of these has a gate and takes a few months to a few years is this very like waterfall mod where the cost of trying things and getting things early is very expensive. But I think to what Matt's saying, if you start to get in a regime where you can have models give you really promising candidates, you can start to make that look a lot more like a loop. Right. It's akin to becoming more agile in software development. But now the next problem is agonists. How do you reliably one shot hitting a...
  4. Latent Space: 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
    ...pitope. You have this equivalent of a content aware fill tool to kind of get your, your binders generated from Chai. And I think to add to that, right? Yeah. This notion of target discovery and hit discovery and optimization where each of these has a gate and takes a few months to a few years is this very like waterfall mod where the cost of trying things and getting things early is very expensive. But I think to what Matt's saying, if you start to get in a regime where you can have models give you really promising candidates, you can start to make that look a lot more like a loop. Right. It's akin to becoming more agile in software development. But now the next problem is agonists. How do you reliably one shot hitting a switch on a cell. Right. Or buy specifics or ADCs. Right. And I think this levels of abstraction that we're going to have to climb with the product a...
Sources
    Latent Space in 3 minutes: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
    Created: August 3rd, 2026 - 14:50 PT
    Script

    Here is The Daily FM summary of the Latent Space that aired on Monday August 3rd. This episode was a deep technical masterclass on inference engineering with Philip Kiely and Ali Taha from Baseten, joining swyx and Vibhu to explain what actually happens after an open model is released and before users experience it as a fast, reliable API. [1]

    The opening question was simple but revealing: what happens when someone sends a 200,000-token prompt? Philip explained that the system first asks whether part of that input has been seen before. If so, cache-aware routing can send the request to a machine that already has some KV cache, avoiding expensive recomputation. If not, the system may split the work between “prefill” GPUs, which process the giant input and generate the first token, and “decode” GPUs, which produce the output tokens. For coding workloads, Baseten may also use speculative decoding, where a smaller model guesses several tokens ahead and the large model verifies them. [2]

    That led into the distinction between shared pay-per-token APIs and dedicated deployments. Ali said dedicated deployments become attractive when traffic is high or specialized, because customers can tune batch sizes, quantization levels, routing, and even train a custom speculative decoder for their own traffic. Philip added that dedicated endpoints also avoid noisy-neighbor problems, like someone else benchmarking a shared API with massive traffic. [3]

    A major theme was that “supporting” a new open model is much more than making it emit one token. Philip said open-source engines like vLLM or SGLang may get basic support quickly, but production readiness requires quantization, calibration, training speculators, testing, routing, and handling new architectural quirks. Ali and Philip gave a striking example from GLM-5.2: Baseten grafted Kimi’s vision encoder onto GLM without changing the language model weights, training only the projector between the “eyes” and the “brain.” Ali said the model learned much better when trained not just to caption images, but to answer detailed questions about them. [4]

    One of the most surprising sections was on failure modes. Models can collapse into repeating the same token, sometimes not because the weights are bad, but because of inference-engine bugs, CUDA kernel race conditions, or differences between clusters and network interconnects. Ali described cases where the same model behaved differently depending on hardware and KV-cache transfer timing. [5]

    The discussion on quantization was especially important. Philip framed quality as fidelity to the original full-precision model. Ali explained that quantization is lossy, but Baseten found that quantizing more layers can sometimes preserve quality better, because errors in different layers cancel each other out. They measure this with KL divergence between logit distributions, not just benchmarks, and claimed this can improve throughput by around 20% while maintaining fidelity. [6]

    The broader takeaway was that inference is still young. Philip said mature fields fight for basis points, while inference optimizations still deliver 20%, 100%, or 200% gains. Stacking NVFP4 quantization, speculative decoding, disaggregated prefill/decode, better kernels, and cache-aware routing can move a model from tens of tokens per second toward several hundred, though the exact gains depend heavily on hardware and traffic. [7]

    The conversation then widened to NVIDIA Dynamo, model parallelism, mega kernels, Rubin, and AI chips. Philip sees Rubin pushing inference toward systems engineering: moving KV cache around clusters, coordinating GPUs, CPUs, and networks, and designing around memory bandwidth. Ali was notably skeptical of mega kernels, arguing that future GPUs are becoming more specialized and that many fused-kernel approaches may not survive in production. [8]

    They also covered video generation, where Ali said open-source video still lags far behind closed models like Veo and Kling. The blocker is attention over enormous numbers of video tokens: five seconds can already mean tens of thousands of tokens, and long videos become brutally expensive. Autoregressive video could enable streaming and longer generation, but today’s quality is poor, while diffusion gives better consistency but struggles to scale to long coherent sequences.

    The episode closed by tying inference back into training. Faster inference helps reinforcement-learning rollouts, while training increasingly has to account for quantization and speculative decoding. Philip predicted continuous loops where deployed models generate traces, get post-trained, A/B tested, and redeployed. Ali gave the memorable example of GLM-5.2 helping profile and write kernels for serving GLM-5.2 itself. The final frontier, they suggested, may be continual learning through persistent KV cache, compacted memory, and models that help optimize the infrastructure they run on. Thank you for listening to Latent Space in 3 minutes from The Daily FM. See you next time! [9]

    Source Evidence
    1. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...s an entirely new optimization problem. In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20% , because the errors introduced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models u...
    2. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200% ; how quantization errors can cancel one another out; why identical weights can behave differently across clus
    3. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...tizing more of the model actually preserved its benchmark quality while increasing throughput by 20% , because the errors introduced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still...
    4. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...cent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20% , because the errors introduced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why infer...
    5. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...ain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200% ; how quantization errors can cancel one another out; why identical weights can behave differently across clus
    6. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...ced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200% ; how quantization errors can cancel one another out; why identical weights can behave d...
    7. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...ning. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200% ; how quantization errors can cancel one another out; why identical weights can behave differently across clus
    8. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...by 20% , because the errors introduced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200% ; how quantization errors can cancel one another out...
    9. Latent Space: The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
      ...s introduced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement , model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200% ; how quantization errors can cancel one another out; why identical weights can...
    Sources
      Latent Space in 3 minutes: Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
      Created: July 28th, 2026 - 08:30 PT
      Script

      Here is The Daily FM summary of the Latent Space that aired on Tuesday July 28th. The episode featured Akshay Nathan from OpenAI, who leads core product engineering for productivity, in a conversation about ChatGPT Work, Codex, and OpenAI’s push to turn coding agents into agents for all knowledge work. [1]

      Akshay framed his career as a long attempt to make software power available to people who do not write code, from no-code and Airtable-style tools to today’s LLM agents. The big product insight, he said, came when OpenAI saw Codex unexpectedly take off among non-developers inside the company. Finance, marketing, and operations people were using it and, more importantly, felt proud of it. They felt like they had gained a new superpower. That convinced OpenAI that the Codex agent harness should not remain just a developer tool. [2]

      A central point was that ChatGPT Work and Codex now share the same underlying agent harness, even though the interfaces differ. Codex still exposes more Git state, diffs, and developer-oriented workflow details. ChatGPT Work hides more of that machinery and focuses on outcomes like spreadsheets, documents, sites, research, and artifacts. Akshay said OpenAI deliberately avoided boxing people into separate products because the boundaries between coding, strategy, design, finance, and operations are blurring. [3]

      The hosts dug into the new interface patterns. Akshay argued that persistent computer environments, plugins, local files, artifacts, and computer use let users delegate much more than a single chat answer. One surprising example was someone using ChatGPT Work to locate a misdelivered package by analyzing a delivery photo and comparing it to nearby apartment listings. Another was OpenAI teams replacing some slide decks and spreadsheets with interactive websites, because sites are more flexible and higher-bandwidth as work artifacts. [4]

      There was also a lot of discussion about models and power-user controls. Akshay’s advice was simple: OpenAI wants the default model configuration to be good enough for most people. Deeper reasoning, Ultra mode, and multi-agent workflows are for unusually complex, exploratory, or parallelizable tasks. The tension, he said, is giving power users control without overwhelming everyone else with toggles.

      Memory was another major theme. Akshay said ChatGPT Work inherits from ChatGPT memory, making it feel like an extension of the user’s existing relationship with the product. He also discussed Chronicle, an experimental way for ChatGPT to learn from computer activity and surface useful context later. The larger idea is that agents become more useful as they accumulate durable context across work and life. [5]

      Some of the most memorable moments were very OpenAI-internal: Akshay described using agents to gather context for performance reviews, not to replace human judgment, but to find contributions a manager might have missed across Slack, docs, and code. He also mentioned an internal scheduled task that reads team activity and generates memes, noting that models are finally becoming genuinely funny.

      The episode closed on how AI changes product development. Akshay said teams can now go from idea to prototype to feedback much faster, which makes ideas, taste, and judgment the bottlenecks. His key warning was not to confuse motion with progress. AI can create more commits, tokens, artifacts, and activity than ever, but teams still need a clear definition of what meaningful progress looks like. Thank you for listening to Latent Space in 3 minutes from The Daily FM. See you next time!

      Source Evidence
      1. Latent Space: Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
        ...ched 10M million users combined (as we cover in the pod, Codex now powers ChatGPT Work, so all ChatGPT Work users are now users of the Codex harness, even if they aren’t traditional engineers) — showing the early innings of what happens when you graduate from coding agents to knowledge work agents: We’ve been calling out how coding agents are “breaking containment” to do everything else this year to power every other part of knowledge work - and it started with the org chart, with a major reorg last month that amounted to two of Codex’s most prominent leaders, Greg and Tibo, taking responsibility over product and ChatGPT specifically, completing a “Superapp” consolidation cycle first discussed in March . With these updates Codex is no longer just a coding tool. In June, OpenAI said knowledge workers already accounting for roughly 20% of Codex’s user base and growing m...
      2. Latent Space: Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
        ...the biggest prize of all — if you can get the agentic interface right. A key trend we have been tracking over at AINews is the absolute explosion in Codex usage this year , with MAU now up >10x from Jan 2026 . Less than two weeks after their July 9th launch , OpenAI said ChatGPT Work and Codex had reached 10M million users combined (as we cover in the pod, Codex now powers ChatGPT Work, so all ChatGPT Work users are now users of the Codex harness, even if they aren’t traditional engineers) — showing the early innings of what happens when you graduate from coding agents to knowledge work agents: We’ve been calling out how coding agents are “breaking containment” to do everything else this year to power every other part of knowledge work - and it started with the org chart, with a major reorg last month that amounted to two of Codex’s most prominent leaders, Greg and Ti...
      3. Latent Space: Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
        ...accounting for roughly 20% of Codex’s user base and growing more than 3x as quickly as developers. A product dedicated for knowledge workers was being pulled out of the Codex team. However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheets for analysis, slide decks for communication, and specialized applications for everything else. ChatGPT Work now enables users to work across every primitive with agents. Instead of opening an application and manually operating its features, the user can describe an outcome and collaborates with an agent that can assemble the tools, context, and arti
      4. Latent Space: Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
        ...tGPT specifically, completing a “Superapp” consolidation cycle first discussed in March . With these updates Codex is no longer just a coding tool. In June, OpenAI said knowledge workers already accounting for roughly 20% of Codex’s user base and growing more than 3x as quickly as developers. A product dedicated for knowledge workers was being pulled out of the Codex team. However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheets for analysis, slide decks for communication, and specialized applications for everything else. ChatGPT Work now enables users to work across every primitive with agents. Instead of opening an application and manually operating its features, the user can describe an outcome and collaborates with an age...
      5. Latent Space: Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
        ...Codex’s user base and growing more than 3x as quickly as developers. A product dedicated for knowledge workers was being pulled out of the Codex team. However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheets for analysis, slide decks for communication, and specialized applications for everything else. ChatGPT Work now enables users to work across every primitive with agents. Instead of opening an application and manually operating its features, the user can describe an outcome and collaborates with an agent that can assemble the tools, context, and arti
      Sources
        Latent Space in 3 minutes: Inside the Model Factory — Eiso Kant, Poolside AI
        Created: July 22nd, 2026 - 22:25 PT
        Script

        Here is The Daily FM summary of the Latent Space that aired on Wednesday July 22nd. This episode went inside Poolside AI with cofounder Eiso Kant, in a very technical but also philosophical conversation about open models, coding agents, and what it takes to build a foundation-model company from scratch. [1]

        Eiso started with his origin story. He said Andrej Karpathy’s 2015 post on recurrent neural nets convinced him overnight that neural networks could learn to write code. He pivoted his earlier startup into “machine learning on code,” spent four or five years and about $12 million pursuing the idea, and ultimately failed because the market was not ready and the team did not fully appreciate how far scaling would go. When ChatGPT arrived, he described it as vindication. [2]

        That history fed into Poolside’s current stance on open models. Eiso said he would rather live in a world with 100 foundation-model companies than five, even if Poolside were one of the five. He distinguished open weights from genuinely open research, arguing that weights alone do not teach others how to reproduce progress. Poolside’s recent technical reports are meant to share more of the machinery, not just benchmark numbers. [3]

        The core of the episode was Poolside’s “model factory.” Eiso argued that model building is 90% engineering: data pipelines, distributed systems, reproducible experiments, reliable training, post-training, and reinforcement learning. Poolside’s team of fewer than 70 researchers reportedly runs 10,000 to 20,000 experiments a month. A major unlock was streaming data directly into training rather than packaging huge datasets in advance. Combined with immutable data and versioned code, that lets them trace experiments down to individual tokens and reproduce old runs. [4]

        A surprising point was how much agents are already involved in Poolside’s own research workflow. Eiso said agents increasingly write code, launch jobs, evaluate results, and modify pipelines, especially in data and synthetic-data work. He framed this as an early glimpse of recursive self-improvement, though humans are still choosing ideas and debugging. [5]

        The new model, Laguna S, was the technical star. It has 118 billion total parameters but only 8 billion active, and Poolside took it from training to launch in about eight weeks. Eiso said its strength is not just raw intelligence but behavior: persistence, verification, backtracking, and refusing to declare victory too early. That led to one of the episode’s key takeaways: smaller models may handle much more knowledge work than people assumed if they are trained to behave like persistent problem-solvers. [6]

        The discussion also looked ahead. Eiso believes reinforcement learning will move earlier into pre-training, because next-token prediction still extracts too little from the web. He criticized the industry’s dependence on distillation and task environments as useful but addictive shortcuts. He also argued against bloated tool-calling systems, calling MCP and traditional tool calls “stupid” in spirit, because future agents should often just write scripts inside containers rather than choose from dozens of predefined tools. [7]

        On safety and policy, Eiso supported democratic oversight as models become more powerful, but warned that premature regulation could lock in an oligopoly of two or three companies. He also emphasized that unilateral safety does not work in a globally competitive race.

        The episode closed with the human side: AI changes engineering productivity by shrinking the time from idea to shipped value, and Eiso said “agency” may become the most important hiring trait. High-agency people, in his view, need shared goals and clear constraints, not micromanagement. Thank you for listening to Latent Space in 3 minutes from The Daily FM. See you next time!

        Source Evidence
        1. Latent Space: Inside the Model Factory — Eiso Kant, Poolside AI
          ...on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1 , that are beating Thinking Machines’ recent release nearly 10 times their size . Poolside’s recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna’s recent technical report on our paper club: From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks , Eiso Kant has spent more than a decade betting that code is the path to AGI . In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with...
        2. Latent Space: Inside the Model Factory — Eiso Kant, Poolside AI
          ...oolside were one of the five. We go deep on Poolside’s Model Factory : the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S , why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web. We also discuss model-harness co-design , Poolside’s path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training,...
        3. Latent Space: Inside the Model Factory — Eiso Kant, Poolside AI
          ...ase nearly 10 times their size . Poolside’s recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna’s recent technical report on our paper club: From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks , Eiso Kant has spent more than a decade betting that code is the path to AGI . In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five. We go deep on Poolside’s Model Factory : the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training,...
        4. Latent Space: Inside the Model Factory — Eiso Kant, Poolside AI
          ...hy Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five. We go deep on Poolside’s Model Factory : the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S , why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web. We also discuss model-harness co-design , Poolsid...
        5. Latent Space: Inside the Model Factory — Eiso Kant, Poolside AI
          ...anies than five even if Poolside were one of the five. We go deep on Poolside’s Model Factory : the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S , why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web. We also discuss model-harness co-design , Poolside’s path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind...
        6. Latent Space: Inside the Model Factory — Eiso Kant, Poolside AI
          ...and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five. We go deep on Poolside’s Model Factory : the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S , why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from t...
        7. Latent Space: Inside the Model Factory — Eiso Kant, Poolside AI
          ...w-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S , why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web. We also discuss model-harness co-design , Poolside’s path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside’s $500 million raise , open-source AI, regulation, NVIDIA and TSMC’s influence , engineering productivity in the age
        Sources
          Latent Space in 3 minutes: Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
          Created: July 21st, 2026 - 12:55 PT
          Script

          Here is The Daily FM summary of the Latent Space that aired on Tuesday July 21st. The episode featured Bo Wang and Ci Chu from Xaira Therapeutics, in a deep conversation about using AI, high-throughput biology, and what they call “virtual cell” models to make drug discovery more predictive and less trial-and-error. [1]

          Xaira’s broader mission, as Chu described it, is to build an AI-native drug discovery company from end to end. They are working on three connected platforms: protein design, drawing on cofounder David Baker’s work; virtual cell models that predict how genes and drugs affect cell biology; and patient representation models that could eventually help identify which patients will respond to which therapies. Wang emphasized that the goal is not just better models, but the right data to power them, connected across target discovery, molecule design, and clinical translation. [2]

          The centerpiece was Xaira’s new model, X-Cell, described as the company’s first virtual cell model. In plain terms, it predicts what happens to a cell when a gene is turned down or knocked out. That matters because many drugs work by inhibiting proteins or pathways, so being able to simulate those effects could help scientists choose better targets before running expensive experiments.

          A major argument in the episode was that causal models need causal data. Chu contrasted observational single-cell datasets, which describe what cells look like, with perturbation datasets, which show what actually happens when you change something. He argued that correlations alone are not enough to learn biology’s cause-and-effect structure. To generate the right training data, Xaira uses perturb-seq, combining CRISPR-based gene perturbations with single-cell RNA sequencing. That lets them knock down genes across huge pools of cells and measure how thousands of genes respond. [3]

          The scale was one of the striking points: Xaira’s dataset includes millions of high-quality cells, multiple genome-wide perturbation campaigns, and many biological contexts. Chu noted that just making this work required major wet-lab engineering, because techniques that work in academic-scale experiments can break down when handling hundreds of millions of cells over long days.

          On the AI side, Wang explained that X-Cell moved away from older autoregressive, GPT-like approaches that require imposing an artificial order on genes. Instead, it uses diffusion language modeling, more like iterative editing, to refine predicted gene-expression states. The model also learns from biological prior knowledge, including literature, protein-protein interaction networks, cancer dependency data, morphology, and prior cell embeddings.

          The most memorable moment was the “wow” result: when the team compared heat maps of predicted gene-expression changes, X-Cell visibly matched ground truth much better than a linear baseline. Even more important, they said the model generalized to settings it had not directly seen, such as activated T cells, held-out cell types, and primary T-cell data from donors, suggesting it may learn transferable biological rules rather than just memorizing screens. [4]

          The hosts also pushed on what comes next. The guests said current virtual cell models are still early and mostly static. Future versions need spatial context, multiple molecular modalities, combinatorial perturbations, organoids, animal models, and eventually patient biology. Chu’s dream bottleneck to remove was high-throughput protein measurement at single-cell scale. Wang’s was even more surprising: sequencing the same living cells over time without killing them, so models can learn true biological dynamics.

          The closing takeaway was that virtual cells are not meant to replace wet labs entirely. They are meant to guide experiments toward harder, more clinically relevant questions, especially where exhaustive testing is impossible. Thank you for listening to Latent Space in 3 minutes from The Daily FM. See you next time!

          Source Evidence
          1. Latent Space: Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
            ...through the podcast is how the lab and experimentation and the real world have probably the biggest impact and have the most relevance to whether something is AI for science or something like B2B SaaS. We're really happy to have in the studio with us today Bo Wang and Si Chu from Xera Therapeutics. At Xera, they're building with a bunch of other people a AI drug discovery platform. They're using high throughput experimentation system to collect very large data sets and then training AI models that can predict the way that your cells in your body will respond to drugs and therapeutics. Really happy to have you. Big fan of your work. Why don't you two introduce yourselves to the listeners?
            
            Speaker B: Hello everyone. My name is Bo Wen. I'm SVP and head of Biomedical AI at Xara Therapeutic. Joined Xara about eight months ago and before that I was associate professor at t...
          2. Latent Space: Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
            ...through the podcast is how the lab and experimentation and the real world have probably the biggest impact and have the most relevance to whether something is AI for science or something like B2B SaaS. We're really happy to have in the studio with us today Bo Wang and Si Chu from Xera Therapeutics. At Xera, they're building with a bunch of other people a AI drug discovery platform. They're using high throughput experimentation system to collect very large data sets and then training AI models that can predict the way that your cells in your body will respond to drugs and therapeutics. Really happy to have you. Big fan of your work. Why don't you two introduce yourselves to the listeners?
            
            Speaker B: Hello everyone. My name is Bo Wen. I'm SVP and head of Biomedical AI at Xara Therapeutic. Joined Xara about eight months ago and before that I was associate professor at t...
          3. Latent Space: Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
            ...RNA therapeutics at Atomic AI and this is the Latent Space AI for Science podcast. One of the themes that has run through the podcast is how the lab and experimentation and the real world have probably the biggest impact and have the most relevance to whether something is AI for science or something like B2B SaaS. We're really happy to have in the studio with us today Bo Wang and Si Chu from Xera Therapeutics. At Xera, they're building with a bunch of other people a AI drug discovery platform. They're using high throughput experimentation system to collect very large data sets and then training AI models that can predict the way that your cells in your body will respond to drugs and therapeutics. Really happy to have you. Big fan of your work. Why don't you two introduce yourselves to the listeners?
            
            Speaker B: Hello everyone. My name is Bo Wen. I'm SVP and head of Bi...
          4. Latent Space: Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
            Speaker A: And what really blew my mind away is when I saw the model make prediction. Just print out the heat map of the genes, correction changes, look at the actual raw data and line up the linear baseline prediction, the ground truth and XL prediction. Altogether it's visually very clear to see that XL prediction is much more similar to ground truth than the linear baseline.
            
            Speaker B: This is a wow moment I was talking about in the beginning.
            
            Speaker A: This is the first time that someone can put together not just one perturbacy but seven genome wide perturbancy campaigns together. Something that jumped out to us biologists right away is that some of the...
          Sources
            Latent Space in 3 minutes: 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
            Created: July 16th, 2026 - 07:06 PT
            Script

            Here is The Daily FM summary of the Latent Space that aired on Thursday July 16th. This episode focused on Lila Sciences, with Andy Beam and Rafa Gómez-Bombarelli explaining a very ambitious thesis: if the internet was the data source that powered large language models, then science itself may be the next internet-scale data generator. [1]

            Andy framed Lila as a bet on the “bitter lesson” of AI: general methods that scale tend to win. His argument was that web data has largely been exhausted, and reinforcement learning has become a way for models to generate and verify their own data in domains like coding and math. Lila wants to do the same thing for science, except the verifier is nature. Their “AI science factories” are automated and semi-automated labs where models propose experiments, instruments or humans execute them, and the resulting data flows back into training. [2]

            A big theme was that Lila is not just a biotech company. Rafa emphasized that they work across life sciences, chemistry, and materials: proteins, RNA, small molecules, catalysts, coatings, quantum dots, polymers, and more. Their belief is that breadth matters. Just as language models benefit from learning code, recipes, and poetry together, Lila thinks a scientific reasoning model can transfer knowledge across domains. One example Rafa gave was that chemistry learned from small-molecule drug discovery helped the system reason about metal-organic frameworks for carbon capture and filtration.

            The most vivid part of the episode was the description of the lab. Andy said the lab of the future should feel less like a room designed for humans and more like a data center: dense, modular, efficient, and always running. Today, Lila connects instruments through a physical transport layer, almost like a computer bus, with plates moving between machines, sometimes on magnetically levitating tracks. But he stressed they are not automation maximalists. If a robot arm makes sense, they use it; if a human hand is faster and cheaper, that can still be “an API call.” [3]

            The hosts pushed on safety, waste, and scientific rigor. Rafa said safety cannot be an afterthought, even if today the platform is used internally by aligned employees. The immediate risks are mostly lab safety issues, like bad chemical combinations or misused instruments, rather than sci-fi malicious behavior. On rigor, both guests argued that AI science has to meet the same standards as human science. One advantage of automation, Andy said, is that experiments can be rerun quickly, with far more metadata captured, down to things like humidity.

            There were also surprising examples of model behavior. Andy described early systems that swore in their chain of thought when asked to redo a 96-well plate map, and models whose “stupid” catalyst suggestions turned out to be the best non-platinum-group electrocatalysts Lila had made. That led into a broader point: distinguishing obviously bad ideas from “Move 37” moments is genuinely hard.

            On the business side, Lila sees the model as the core asset, while the lab is the token generator. Andy compared the platform to a “Claude Code for science,” letting partners run virtual startups without building their own lab. He described an in vivo CAR-T effort where a tiny internal team combined binder design, lipid nanoparticle formulation, and mRNA design, reaching strong non-human-primate data in about six months. Lila does not want to become a single-asset drug company; instead, it wants many partners building on the platform.

            The closing takeaway was that Lila is trying to move science up the abstraction ladder. Scientists should spend less time manually translating ideas into protocols and more time asking better questions. The bottleneck, in Lila’s view, is not just compute or data, but the ability to turn experiments into scalable, verified training signals. Thank you for listening to Latent Space in 3 minutes from The Daily FM. See you next time! [4]

            Source Evidence
            1. Latent Space: 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
              ...cale dataset coming from?
              
              Speaker A: You know, people normally talk about different scaling axes. You have compute, you have data, and for science, data is not necessarily an infinite resource. And your point is that we now want to add a new scaling axis for data.
              
              Speaker B: We think that like the lab of the future should feel like a data center, rows of server racks as densely packed as possible and also as energy efficient as possible and things like that.
              
              Speaker A: Welcome to Latent Space Science. I'm Brandon. I'm here with my co-host RJ. Today we have Rafa Gomez-Bombarelli and Andy Beam from Lila Science. We'll just start off and let you introduce yourself.
              
              Speaker B: Yeah, thanks for having us on the podcast. Like, you know, longtime listener, first-time caller. Excited to be here. I'm Andy. I'm the Chief Technology Officer at Lila. I've been an AI researche...
            2. Latent Space: 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
              Speaker A: So not just TechBio, what do you do in terms of science?
              
              Speaker B: We are all in on the bitter lesson and scale. We think that methods that scale and that are general beat those that are not. You know, as Ilya said at NeurIPS last year, we have but one internet. It's the fossil fuel we fracked. We got every ounce of data that we could out of the internet, but it's gone. And so the question in AI is like, where is the next internet-scale dataset coming from?
              
              Speaker A: You know, people normally talk about different scaling axes. You have compute, you have data, and for science, data is not necessarily an infinite resource. And your point is that we now want...
            3. Latent Space: 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
              ...at we could out of the internet, but it's gone. And so the question in AI is like, where is the next internet-scale dataset coming from?
              
              Speaker A: You know, people normally talk about different scaling axes. You have compute, you have data, and for science, data is not necessarily an infinite resource. And your point is that we now want to add a new scaling axis for data.
              
              Speaker B: We think that like the lab of the future should feel like a data center, rows of server racks as densely packed as possible and also as energy efficient as possible and things like that.
              
              Speaker A: Welcome to Latent Space Science. I'm Brandon. I'm here with my co-host RJ. Today we have Rafa Gomez-Bombarelli and Andy Beam from Lila Science. We'll just start off and let you introduce yourself.
              
              Speaker B: Yeah, thanks for having us on the podcast. Like, you know, longtime listener, first...
            4. Latent Space: 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
              ...cale dataset coming from?
              
              Speaker A: You know, people normally talk about different scaling axes. You have compute, you have data, and for science, data is not necessarily an infinite resource. And your point is that we now want to add a new scaling axis for data.
              
              Speaker B: We think that like the lab of the future should feel like a data center, rows of server racks as densely packed as possible and also as energy efficient as possible and things like that.
              
              Speaker A: Welcome to Latent Space Science. I'm Brandon. I'm here with my co-host RJ. Today we have Rafa Gomez-Bombarelli and Andy Beam from Lila Science. We'll just start off and let you introduce yourself.
              
              Speaker B: Yeah, thanks for having us on the podcast. Like, you know, longtime listener, first-time caller. Excited to be here. I'm Andy. I'm the Chief Technology Officer at Lila. I've been an AI researche...
            Sources
              Latent Space in 3 minutes: Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
              Created: July 9th, 2026 - 11:05 PT
              Script

              Here is The Daily FM summary of the Latent Space that aired on Wednesday July 8th. The episode featured Modal CTO Akshat Bubna, alongside cofounder Vibhu, in a wide-ranging conversation about how AI infrastructure is changing as agents, inference, and bursty GPU workloads become central to modern software. [1]

              Akshat began by revisiting Modal’s origin story. The company did not start as a GPU inference platform, but as an attempt to build a better runtime than Kubernetes for compute-heavy, bursty workloads. The key idea was that developers should not have to manage piles of YAML or Kubernetes configuration just to run jobs with custom images, accelerators, or fast scaling. Modal’s answer was “self-provisioning workloads”: infrastructure requirements live next to the code, often as decorators, so the system can spin up the right compute automatically. [2]

              A major theme was the shift from developer experience to what Akshat called agent experience. Modal has apparently reorganized its SDK thinking around how AI agents use infrastructure. His argument was simple: if it is painful for humans to read and modify hundreds of Kubernetes files, it is also painful for agents. Typed, colocated infrastructure definitions are easier for coding agents to modify, test, and observe. The hosts pushed on whether this still matters if humans are not reading code as much anymore, and Akshat said observability may now be even more important: agents can change code, but humans still need dashboards, logs, and judgment to understand what happened.

              The conversation then moved into Modal’s current identity. Akshat described it as a cloud platform with primitives built from scratch for AI applications, covering inference, training, batch processing, and sandboxes. He emphasized that Modal is not trying to replace always-on web hosting. Its sweet spot is specialized compute that scales up and down fast, across GPUs, CPUs, regions, and unusual workloads.

              One of the most interesting sections covered sandboxes. Modal built sandbox APIs as early as 2023, before coding agents had fully taken off, and even used an early “small developer” agent loop as an example. Back then, models tended to diverge after a handful of iterations. In hindsight, the hosts joked, the winning move would have been to collect all those failures, build benchmarks and RL environments, and turn them into a billion-dollar agent company. [3]

              Akshat also explained that Modal’s biggest use case today is elastic inference for custom models, especially in audio, video, robotics, and computational biology. Companies like Suno and Runway may train models elsewhere but use Modal to deploy and autoscale them. Modal has gone deep on cold starts, GPU snapshotting, and regional autoscaling, because production inference is not just “find a GPU and run a model”; it involves tail latency, reliability, request delivery, and fast scaling across regions. [4]

              A technical highlight was speculative decoding. Akshat explained that a smaller draft model predicts tokens ahead, while the larger model verifies them. If many drafted tokens are accepted, inference can become two to four times faster without quality loss. Modal’s open-source DFlash work uses block-based speculation, and its new Auto Endpoints aim to make optimized open-model serving easier while still letting users eject into full code when they need customization.

              The episode also touched on distributed training, RDMA networking, private IPv6 overlay networks, and why Modal runs across 17 cloud and neocloud providers rather than owning data centers. Akshat framed Modal as a “supercloud” software layer, with its own reliability and scheduling layer on top of many providers.

              Near the end, the discussion broadened to AI infrastructure trends: agent sandboxes, CI for coding agents, continual learning, auto-research, robotics, drug discovery, and whether future video generation may be orchestrated by agents rather than single video models. The final takeaway was that Modal’s bet on developer experience has evolved naturally into agent experience, and that the infrastructure primitives that once seemed niche now look central to the next generation of AI products. Thank you for listening to Latent Space in 3 minutes from The Daily FM. See you tomorrow!

              Source Evidence
              1. Latent Space: Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
                Speaker A: We're here with Akshat of Modo, CTO of Modo, together with Vibhu. Congrats on your Series C.
                
                Speaker B: Thank you.
                
                Speaker A: Your party yesterday was amazing.
                
                Speaker B: Yeah.
                
                Speaker A: All the photos and all the swag.
                
                Speaker B: We, we had a bunch of art installations, which is kind of fun seeing like our products on pedestals next to like Rodin.
                
                Speaker A: Very nice. Very nice. When you started, it was not the GPU inference company. I mean, maybe it was in your mind. Take us back to the origin story.
                
                Speaker B: I actually first met Eric, who's the CEO, through an investor. And back then, Eric was already thi...
              2. Latent Space: Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
                ....
                
                Speaker A: All the photos and all the swag.
                
                Speaker B: We, we had a bunch of art installations, which is kind of fun seeing like our products on pedestals next to like Rodin.
                
                Speaker A: Very nice. Very nice. When you started, it was not the GPU inference company. I mean, maybe it was in your mind. Take us back to the origin story.
                
                Speaker B: I actually first met Eric, who's the CEO, through an investor. And back then, Eric was already thinking about building a new kind of runtime. And he got there thinking through why are workflow orchestration products so hard to use? It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and custom images, and it has a terrible developer experience.
                
                Speaker A: And I'll inject for listeners who are new, we interviewed Eric 2 years ago, and there's a bit more of the story th...
              3. Latent Space: Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
                ...r. And back then, Eric was already thinking about building a new kind of runtime. And he got there thinking through why are workflow orchestration products so hard to use? It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and custom images, and it has a terrible developer experience.
                
                Speaker A: And I'll inject for listeners who are new, we interviewed Eric 2 years ago, and there's a bit more of the story there from Spotify and all those things. And I actually came across Eric through Data Council because he did that talk on the sort of serverless container stack that you guys did, which is like That was my first, like, okay, I need to take models very seriously moment. But it was still very unclear, like, do I actually need all this for just my data pipelines?
                
                Speaker B: Yeah, I mean, initially what we were...
              4. Latent Space: Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
                ...s already thinking about building a new kind of runtime. And he got there thinking through why are workflow orchestration products so hard to use? It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and custom images, and it has a terrible developer experience.
                
                Speaker A: And I'll inject for listeners who are new, we interviewed Eric 2 years ago, and there's a bit more of the story there from Spotify and all those things. And I actually came across Eric through Data Council because he did that talk on the sort of serverless container stack that you guys did, which is like That was my first, like, okay, I need to take models very seriously moment. But it was still very unclear, like, do I actually need all this for just my data pipelines?
                
                Speaker B: Yeah, I mean, initially what we were thinking about was if we...
              Sources

                <- Back to library