[
  {
    "start": 0.0,
    "end": 6.18,
    "text": "we start seeing people adding depth, for example to these models or a sense of touch."
  },
  {
    "start": 7.4,
    "end": 7.5,
    "text": "Okay?"
  },
  {
    "start": 7.56,
    "end": 12.92,
    "text": "We have this hardware constraint that you typically only like say one or two GPUs on our robot."
  },
  {
    "start": 13.5,
    "end": 15.78,
    "text": "there are people who actually said were not solving the problem."
  },
  {
    "start": 15.84,
    "end": 21.62,
    "text": "right now let's put a big GPU cluster next to it and then we had five hundred millisecond latency in that."
  },
  {
    "start": 21.8,
    "end": 24.06,
    "text": "still okay if he has all other layers in place"
  },
  {
    "start": 24.58,
    "end": 24.8,
    "text": "how do"
  },
  {
    "start": 25.9,
    "end": 28.06,
    "text": "decide what to keep and how to compress."
  },
  {
    "start": 28.34,
    "end": 30.42,
    "text": "I mean, honestly through learning right?"
  },
  {
    "start": 30.68,
    "end": 36.8,
    "text": "So you basically learn the architecture such that the compression mechanism is just very efficient."
  },
  {
    "start": 40.4,
    "end": 42.18,
    "text": "Hello and welcome back to Rob Talk!"
  },
  {
    "start": 42.46,
    "end": 45.08,
    "text": "The podcast about physical AI and robotics."
  },
  {
    "start": 45.78,
    "end": 60.96,
    "text": "And today we're gonna talk about the latest and greatest About what's going on What is evolving in the context of physical AI and robotics in two thousand twenty six."
  },
  {
    "start": 61.64,
    "end": 66.56,
    "text": "And we cover things like world models, what is that?"
  },
  {
    "start": 66.68,
    "end": 71.54,
    "text": "Is it the thing where the answer's always forty-two or something else?"
  },
  {
    "start": 71.98,
    "end": 78.54,
    "text": "for that I have Felix with me, Felix Frank our staff engineer on the robot intelligence team at Robco."
  },
  {
    "start": 79.78,
    "end": 80.38,
    "text": "Yeah, Felix!"
  },
  {
    "start": 82.02,
    "end": 88.16,
    "text": "What is the gist of Two Thousand Twenty Six?"
  },
  {
    "start": 88.3,
    "end": 90.32,
    "text": "The Autonomous Robotics Podcast."
  },
  {
    "start": 90.88,
    "end": 94.24,
    "text": "Physical AI, no theory just reality."
  },
  {
    "start": 98.96,
    "end": 104.2,
    "text": "I think we can start at a high level and then go down into more depth on some of the topics."
  },
  {
    "start": 104.84,
    "end": 117.72,
    "text": "but i guess there's few trends where seeing people switching from VLAs or vision language action models to world model video models And will talk about that."
  },
  {
    "start": 118.62,
    "end": 134.0,
    "text": "But theres also other things like having policies that can deal with longer horizon tasks, so more high level abstract tasks which have to be completed by adding memory or things like."
  },
  {
    "start": 134.2,
    "end": 139.22,
    "text": "So I think there's a lot of really exciting developments in the space right now."
  },
  {
    "start": 139.84,
    "end": 145.4,
    "text": "There is a lot going on and yeah, twenty-twenty six is definitely big year in physical AI."
  },
  {
    "start": 145.5,
    "end": 146.72,
    "text": "All right, yeah."
  },
  {
    "start": 147.4,
    "end": 149.74,
    "text": "We've seen this huge jumps over the course of a year."
  },
  {
    "start": 150.06,
    "end": 151.54,
    "text": "so let's start to decode it little bit."
  },
  {
    "start": 152.74,
    "end": 161.98,
    "text": "when we look back people have been controlling robots for decades and since few years ago also with learning."
  },
  {
    "start": 162.82,
    "end": 170.08,
    "text": "but what was thing that really changed things in these two thousand twenty three, two thousand forty four time frame?"
  },
  {
    "start": 170.58,
    "end": 182.7,
    "text": "I think you can even start earlier like end of twenty-twenty where previously we had a lot of small, special models for one specific task."
  },
  {
    "start": 182.8,
    "end": 201.44,
    "text": "And then with the language models becoming so good and so prominent I mean...I think we all remember chat GPT like coming to the public and people seeing oh there's actually value in big models and a lot data in scaling this entire process right?"
  },
  {
    "start": 202.04,
    "end": 205.0,
    "text": "Then robotics followed actually quickly after."
  },
  {
    "start": 205.88,
    "end": 213.18,
    "text": "So there's works like RT-One, RT-Two, Palm E where the idea really was okay."
  },
  {
    "start": 213.26,
    "end": 230.82,
    "text": "we can kind of take the same architecture that we have in language models which is transformers and they don't just have to be used for language but they're essentially very good at any type of sequential data and robotics data in and off itself."
  },
  {
    "start": 233.56,
    "end": 236.3,
    "text": "You have joint measurements, you have images."
  },
  {
    "start": 237.52,
    "end": 250.88,
    "text": "Maybe force measurement like descriptions of what is going on and all that is a sequence And basically also try to model and solve these problems with transformers."
  },
  {
    "start": 251.4,
    "end": 264.18,
    "text": "So in twenty-twenty two and twenty three were really when people started saying okay What if we take this paradigm of gathering data making bigger models put that to robotics."
  },
  {
    "start": 264.3,
    "end": 277.62,
    "text": "And, and that's kind of where the origin off vision language action models is what we these days call like well let say The Gold Standard Of What We Were Doing In The Past Years."
  },
  {
    "start": 277.98,
    "end": 282.58,
    "text": "um I think so if i structured a bit there are two directions."
  },
  {
    "start": 283.2,
    "end": 291.76,
    "text": "One Direction Is What I Just Talked About Which Is This Like Larger Models More Data Yeah, learn from what the language models were doing."
  },
  {
    "start": 291.88,
    "end": 298.56,
    "text": "And then the other big paradigm which I think we have to talk about is... Is what we call these days imitation learning?"
  },
  {
    "start": 299.06,
    "end": 302.14,
    "text": "Which has actually been around robotics for a really long time."
  },
  {
    "start": 303.02,
    "end": 308.48,
    "text": "but in like twenty-twenty three people started to really push that forward."
  },
  {
    "start": 308.54,
    "end": 323.1,
    "text": "so they where papers form DeepMind called Aloha and Liege For Example Where They Really Said Okay How Far Can We Push The idea of I show the robot a certain task, it can be dexterous."
  },
  {
    "start": 323.2,
    "end": 327.14,
    "text": "It can be a certain pick-and-place tasks that can have soft objects or something like this."
  },
  {
    "start": 327.92,
    "end": 329.02,
    "text": "and how far Can i push?"
  },
  {
    "start": 329.08,
    "end": 331.8,
    "text": "This just by sheer number Of demonstration."
  },
  {
    "start": 331.9,
    "end": 353.68,
    "text": "so i remember reading that paper back then And i think they had um that thirty five operators for Like a bimanual setup So two robot arms and they were collecting data over eight months Um and then using basically what is an early version of like a diffusion policy model to basically try to learn this."
  },
  {
    "start": 354.28,
    "end": 357.4,
    "text": "So if I summarize, right?"
  },
  {
    "start": 357.44,
    "end": 359.56,
    "text": "There's these two directions."
  },
  {
    "start": 359.72,
    "end": 367.68,
    "text": "one is okay we want really go into learning from data and scale up our data collection effort."
  },
  {
    "start": 367.92,
    "end": 370.4,
    "text": "the other one then large models."
  },
  {
    "start": 371.72,
    "end": 382.64,
    "text": "there also an important step which you quickly realize that the amount of robotics data that we have is just limited."
  },
  {
    "start": 383.36,
    "end": 385.0,
    "text": "And very costly to collect, right?"
  },
  {
    "start": 385.48,
    "end": 394.38,
    "text": "We know that collecting tele-operation data is quite slow and then... That's also where vision language action models come from."
  },
  {
    "start": 394.46,
    "end": 395.32,
    "text": "The idea was okay."
  },
  {
    "start": 395.38,
    "end": 403.5,
    "text": "there are other data modalities out there like text or images which we have many magnitudes more Right!"
  },
  {
    "start": 404.24,
    "end": 412.42,
    "text": "Then we had really good examples on how these models learn amazing capabilities with JetGPT and so on."
  },
  {
    "start": 413.9,
    "end": 421.16,
    "text": "And you basically say, okay how can I take this knowledge which is available in a pre-trained model right?"
  },
  {
    "start": 421.36,
    "end": 425.16,
    "text": "Then combine it with the imitation learning paradigm and robot data."
  },
  {
    "start": 425.54,
    "end": 429.34,
    "text": "that's kind of what where vision language action models come from."
  },
  {
    "start": 430.96,
    "end": 432.66,
    "text": "So different developments."
  },
  {
    "start": 434.0,
    "end": 438.56,
    "text": "first we had this jet GPT moment four years ago, right?"
  },
  {
    "start": 438.72,
    "end": 439.1,
    "text": "Pretty much."
  },
  {
    "start": 439.98,
    "end": 442.76,
    "text": "Incredibly long time ago has changed so much."
  },
  {
    "start": 443.26,
    "end": 455.28,
    "text": "So we suddenly saw that more data helps in this whole language space and it starts to spill over And from a vision world there wasn't really that much data."
  },
  {
    "start": 455.6,
    "end": 461.02,
    "text": "There's the bootstrapping problem but their labs also combined their datasets."
  },
  {
    "start": 461.32,
    "end": 467.56,
    "text": "Suddenly We had something like ten thousand or fifty thousand hours somehow with some embodiment."
  },
  {
    "start": 467.68,
    "end": 474.28,
    "text": "Okay, so now this is put together and what was the quality that you could achieve?"
  },
  {
    "start": 474.34,
    "end": 477.12,
    "text": "This was the first time where we had these really large models"
  },
  {
    "start": 477.64,
    "end": 477.88,
    "text": "right?"
  },
  {
    "start": 477.92,
    "end": 477.98,
    "text": "yeah?"
  },
  {
    "start": 478.02,
    "end": 488.54,
    "text": "So I mean i think back then when people were looking at us okay how good are these models in like zero short capabilities?"
  },
  {
    "start": 488.6,
    "end": 494.5,
    "text": "or You basically have your let's say fifty thousand hours of like diverse robot data."
  },
  {
    "start": 494.72,
    "end": 502.48,
    "text": "You combine that with a vision language backbone, which is basically very good at understanding text and understanding certain task descriptions right?"
  },
  {
    "start": 502.96,
    "end": 520.2,
    "text": "And then you see how good this combination in solving tasks... I mean the results were impressive but not at the level of where you would need them to be actually usable, right?"
  },
  {
    "start": 520.26,
    "end": 522.98,
    "text": "So like certain lab benchmarks."
  },
  {
    "start": 523.62,
    "end": 531.32,
    "text": "Certain standard tasks these models could basically do zero shot with and all thirty forty fifty percent success rate."
  },
  {
    "start": 531.82,
    "end": 538.96,
    "text": "but um... But he didn't really have models that went all the way to ninety ninety plus percent success rates."
  },
  {
    "start": 539.54,
    "end": 541.38,
    "text": "so I think there was."
  },
  {
    "start": 543.12,
    "end": 545.74,
    "text": "There were basically a few things that we needed to figure out."
  },
  {
    "start": 547.66,
    "end": 560.46,
    "text": "And then at the same time, I guess there's always this distinction between having a model is like very good generically on a range of tasks and models that really specialize one specific task you want solve."
  },
  {
    "start": 561.92,
    "end": 569.74,
    "text": "Okay so... We have these vision backbones which has been trained for long times now with language part coming in."
  },
  {
    "start": 570.76,
    "end": 575.52,
    "text": "but So now you can deal with this in several different dimensions."
  },
  {
    "start": 575.84,
    "end": 581.28,
    "text": "One is adding more data and the data has been growing, right?"
  },
  {
    "start": 583.02,
    "end": 590.38,
    "text": "We are seeing labs that talking about millions of hours somehow annotated data... ...in a video space."
  },
  {
    "start": 591.86,
    "end": 593.12,
    "text": "The other part is modeling."
  },
  {
    "start": 595.52,
    "end": 597.14,
    "text": "a model is of function, right?"
  },
  {
    "start": 597.18,
    "end": 598.78,
    "text": "It has inputs and outputs."
  },
  {
    "start": 599.18,
    "end": 601.92,
    "text": "So what's the input to such a VLA?"
  },
  {
    "start": 604.64,
    "end": 625.64,
    "text": "I mean usually the input for such a vla at that most basic level like a snapshot of current state which means basically camera view on everything it sees in most setups or more industrial setups."
  },
  {
    "start": 625.68,
    "end": 628.0,
    "text": "You usually have a bimanual arm, right?"
  },
  {
    "start": 628.04,
    "end": 638.08,
    "text": "So you have two robot arms and then both of these robot arms have cameras on the wrists And then you usually have one static camera that is fixed as a head camera or like static to the scene."
  },
  {
    "start": 638.12,
    "end": 646.02,
    "text": "in general The input to the model at that point would be the joint readings Of all-of all the joints in both robots."
  },
  {
    "start": 646.08,
    "end": 653.08,
    "text": "so an idea of what configuration Is the robot currently in plus camera images?"
  },
  {
    "start": 653.3,
    "end": 655.24,
    "text": "uh, yeah."
  },
  {
    "start": 655.3,
    "end": 657.5,
    "text": "Of the current like what the camera seeing?"
  },
  {
    "start": 657.54,
    "end": 658.12,
    "text": "Uh."
  },
  {
    "start": 658.44,
    "end": 666.08,
    "text": "and then usually on top of that you have what we call a language prompt which is exactly what-what do would expect when you're dealing with language models?"
  },
  {
    "start": 666.12,
    "end": 673.36,
    "text": "so it's basically an input That is supposed to Give The model An idea of What It's Supposed To Be Doing."
  },
  {
    "start": 673.46,
    "end": 675.62,
    "text": "So Its Kind Of Like A Goal Description Mm-hmm!"
  },
  {
    "start": 675.68,
    "end": 683.96,
    "text": "So state space Is Such An Important Concept In Robotics Right Given That You Cannot input, everything that you know about the world at all times."
  },
  {
    "start": 684.02,
    "end": 688.64,
    "text": "You have to somehow limit yourself in what actually put it and want to ignore."
  },
  {
    "start": 689.1,
    "end": 698.44,
    "text": "so It seems like having just this stated one particular point of time is trying To find out What a conversation's about by hearing a single sentence?"
  },
  {
    "start": 698.84,
    "end": 699.22,
    "text": "Yeah Right."
  },
  {
    "start": 700.58,
    "end": 706.42,
    "text": "So let's talk A little bit About what The limits Is And what people Have been Trying to do."
  },
  {
    "start": 706.52,
    "end": 707.2,
    "text": "with That"
  },
  {
    "start": 707.24,
    "end": 708.8,
    "text": "I mean there's really good Point."
  },
  {
    "start": 708.86,
    "end": 726.72,
    "text": "because So, I mean in like sequence modeling or robotics we call this the Markov assumption that you can essentially from your state at a current point of time perfectly predict where are and what is going."
  },
  {
    "start": 726.92,
    "end": 735.96,
    "text": "And then it's just well defined problem but basically means to carefully design this observation."
  },
  {
    "start": 736.02,
    "end": 736.96,
    "text": "which model is getting?"
  },
  {
    "start": 737.42,
    "end": 744.9,
    "text": "because you can easily run into issues where if, I mean... You just imagine going through a labyrinth."
  },
  {
    "start": 745.02,
    "end": 746.92,
    "text": "Where all the walls look the same right?"
  },
  {
    "start": 747.22,
    "end": 759.7,
    "text": "If you get one snapshot and don't have history of how you moved before it's basically impossible to like keep track of where you are and where your going."
  },
  {
    "start": 760.06,
    "end": 761.84,
    "text": "And it's the same thing with a conversation, right?"
  },
  {
    "start": 762.14,
    "end": 767.52,
    "text": "If there is one hour conversation then you hear one sentence out of that basically missing alot context."
  },
  {
    "start": 768.08,
    "end": 781.16,
    "text": "so... There're different ways how can fix this let say or different way how people try to approach these issues having some sort of memory or history."
  },
  {
    "start": 782.88,
    "end": 792.64,
    "text": "The easiest usually just Getting the model some form of velocity information, which usually means you just add multiple frames off."
  },
  {
    "start": 792.72,
    "end": 793.94,
    "text": "The same observation space."
  },
  {
    "start": 794.04,
    "end": 796.88,
    "text": "so instead of Just seeing where You currently are?"
  },
  {
    "start": 797.08,
    "end": 801.1,
    "text": "You basically also see oh by the way this is Where have been half a second ago."
  },
  {
    "start": 801.28,
    "end": 804.96,
    "text": "This Is where I've Been A Second ago and you can extend that quite a bit."
  },
  {
    "start": 805.92,
    "end": 810.0,
    "text": "And That Way the the Model Basically has an idea Of all."
  },
  {
    "start": 810.72,
    "end": 814.32,
    "text": "This is where I have been going, right?"
  },
  {
    "start": 814.76,
    "end": 820.04,
    "text": "And then together with this goal description of where i want to be you can maybe figure out okay."
  },
  {
    "start": 820.12,
    "end": 821.1,
    "text": "Where do we go next?"
  },
  {
    "start": 821.16,
    "end": 829.96,
    "text": "because that's ultimately what we want these models to basically output a next action which is bringing us closer to the goal."
  },
  {
    "start": 832.2,
    "end": 836.52,
    "text": "so one option is kind of engineering your observation space."
  },
  {
    "start": 839.1,
    "end": 847.34,
    "text": "Another option, and that's kind of more recent is you basically add memory directly to the model."
  },
  {
    "start": 847.84,
    "end": 849.48,
    "text": "And there are different ways how you could do it."
  },
  {
    "start": 850.46,
    "end": 864.9,
    "text": "There's a whole line of research which basically tries to add memory either in the form some sort of database or like past experiences That can efficiently query while doing this task."
  },
  {
    "start": 865.32,
    "end": 866.92,
    "text": "So you have like a reference point."
  },
  {
    "start": 867.02,
    "end": 869.16,
    "text": "Oh, this is something I've done yesterday."
  },
  {
    "start": 869.3,
    "end": 871.3,
    "text": "and then what was the experience there?"
  },
  {
    "start": 872.64,
    "end": 877.1,
    "text": "Or another option as... You can also imagine it's an internal monologue."
  },
  {
    "start": 877.48,
    "end": 887.18,
    "text": "so we have some sort of higher level agent that usually a VLM or vision language model And basically orchestrating supervising execution."
  },
  {
    "start": 889.28,
    "end": 897.02,
    "text": "Yeah, really having an inner monologue and like from time to time checking up Oh what is my vision language action model?"
  },
  {
    "start": 897.06,
    "end": 901.12,
    "text": "So the let's say lower level part doing And then also correcting it."
  },
  {
    "start": 902.56,
    "end": 917.44,
    "text": "I mean i think in The context of Robco and of industrial task right...I Think we should also talk about the most easiest way Of doing that which is essentially just as said before That part of the observation space Is adding a prompt or goal description."
  },
  {
    "start": 918.28,
    "end": 926.82,
    "text": "And especially in industrial tasks, when you don't deal with a lot of variety and your task."
  },
  {
    "start": 927.0,
    "end": 935.54,
    "text": "You know pretty well ahead-of time what is the usual sequence that has to be solved?"
  },
  {
    "start": 936.24,
    "end": 940.6,
    "text": "Then also do this in more classical engineering way."
  },
  {
    "start": 940.68,
    "end": 949.12,
    "text": "so basically have some state machine around it which orchestrates switching this prompt, so the language instruction once you complete a sub-task."
  },
  {
    "start": 949.18,
    "end": 950.8,
    "text": "And then you need some system to figure out."
  },
  {
    "start": 951.26,
    "end": 959.56,
    "text": "okay now certain subtasks is completed but that's kind of easiest and most engineering way But overall there are like a variety different ways."
  },
  {
    "start": 959.76,
    "end": 960.62,
    "text": "how can tackle it?"
  },
  {
    "start": 961.42,
    "end": 965.38,
    "text": "So if I want summarize maybe take another example."
  },
  {
    "start": 965.84,
    "end": 977.38,
    "text": "ten years ago when we were using Google Translate The limiting factor was state space like what the next word should be."
  },
  {
    "start": 978.84,
    "end": 983.54,
    "text": "You only had enough data to look maybe one or two words behind, right?"
  },
  {
    "start": 983.6,
    "end": 993.76,
    "text": "And so I think this whole thing with transformers and chat GPT show that you can basically have an infinite context if you know what to attend too."
  },
  {
    "start": 993.98,
    "end": 1003.82,
    "text": "The motivation came from machine translation because of language A in a different order than we need where to look towards."
  },
  {
    "start": 1004.34,
    "end": 1016.66,
    "text": "And I think this is here sort of comparable only that you also have a history of what was happening in the near past, which could give you some information on what to do in the future?"
  },
  {
    "start": 1016.9,
    "end": 1027.839,
    "text": "Yeah and so it seems like we've been starting with most limited context that you can possibly use but now were trying to expand."
  },
  {
    "start": 1028.04,
    "end": 1028.04,
    "text": "I"
  },
  {
    "start": 1031.339,
    "end": 1035.04,
    "text": "think an important point is also you're completely right."
  },
  {
    "start": 1035.54,
    "end": 1038.099,
    "text": "The context length of Transformers, it's amazing!"
  },
  {
    "start": 1038.48,
    "end": 1040.7,
    "text": "You can have the entire book in there pretty much by now."
  },
  {
    "start": 1041.359,
    "end": 1065.72,
    "text": "and the tricky thing with robotics obviously always that... ...you need to handle this whole context at a very fast pace because we are controlling our robots at something like fifty hertz usually or thirty hertz which basically means.. ..you need new action input thirty to fifty two hundred times per second, which gives you a very small amount of computation time."
  },
  {
    "start": 1066.26,
    "end": 1075.28,
    "text": "And that's also why this idea of okay I have somehow orchestrate the state space well or how am i efficient about my memory?"
  },
  {
    "start": 1075.58,
    "end": 1076.98,
    "text": "Why is it so important in robotics?"
  },
  {
    "start": 1077.04,
    "end": 1091.74,
    "text": "because You don't have an huge cluster right That can basically just go through all those numbers at A quick pace and where it doesn't matter whether point three seconds later, right?"
  },
  {
    "start": 1092.18,
    "end": 1093.36,
    "text": "But but in robotics."
  },
  {
    "start": 1093.46,
    "end": 1097.06,
    "text": "I mean you really need this high frequency interaction"
  },
  {
    "start": 1097.86,
    "end": 1103.1,
    "text": "and so having more context means usually bigger models more compute lower latency."
  },
  {
    "start": 1103.44,
    "end": 1104.64,
    "text": "there's some tricks around that."
  },
  {
    "start": 1104.74,
    "end": 1121.04,
    "text": "Right i mean first of all we were talking We have talked In the previous episode about The different layers That You Need Something Like One millisecond Maybe Latency And Then You Have These These Models That Were Talking About Where Sixty Hertz Would Be Sixteen Milliseconds But then they are usually slower, right?"
  },
  {
    "start": 1121.34,
    "end": 1123.06,
    "text": "So there's some tricks around that as well."
  },
  {
    "start": 1123.46,
    "end": 1136.24,
    "text": "Yeah I mean you usually have some layered architecture and i mean we've talked about this in the past episode And I think you've talked a bunch of past episodes where You have at very lowest layer Usually quiet deterministic."
  },
  {
    "start": 1136.32,
    "end": 1153.1,
    "text": "so your classical robot controller which runs at kilohertz speed Then if actually go all the way to motor control it is even faster than Kilohertz Speed have layers of models or controllers that basically run at slower frequencies."
  },
  {
    "start": 1153.42,
    "end": 1164.88,
    "text": "So your VLA, your vision language action model would usually run somewhere around thirty to a hundred hertz maybe and like deal with not just the lowest level."
  },
  {
    "start": 1165.14,
    "end": 1170.08,
    "text": "okay I have my joint in this exact position but think about a bit longer term."
  },
  {
    "start": 1170.18,
    "end": 1171.86,
    "text": "Okay, I'm here i want to be there."
  },
  {
    "start": 1172.02,
    "end": 1173.26,
    "text": "how do have to move and so on?"
  },
  {
    "start": 1173.74,
    "end": 1183.8,
    "text": "And then this is that can be an even higher level orchestrator which would just stand for example vision language model with runs at more like one Hertz or maybe once the second checks in on."
  },
  {
    "start": 1184.12,
    "end": 1187.56,
    "text": "okay I am still on my way from the kitchen to the bedroom right."
  },
  {
    "start": 1187.96,
    "end": 1188.6,
    "text": "continue all."
  },
  {
    "start": 1188.88,
    "end": 1190.24,
    "text": "now I've arrived at the bad room."
  },
  {
    "start": 1190.32,
    "end": 1204.18,
    "text": "so no my task changes something but it's basically layer of of models that run at different frequencies and have different tasks from lowest level control to highest level orchestration."
  },
  {
    "start": 1204.68,
    "end": 1205.14,
    "text": "And reasoning?"
  },
  {
    "start": 1206.34,
    "end": 1212.56,
    "text": "Okay, we had this hardware constraint you typically only one or two GPUs on a robot."
  },
  {
    "start": 1213.78,
    "end": 1225.64,
    "text": "There are people who say they're not solving the problem right now let's put a big GPU cluster next And then we have five hundred millisecond latency, and that's still okay if you have all the other layers in place."
  },
  {
    "start": 1226.46,
    "end": 1229.56,
    "text": "But obviously would like it to be as self-contained is possible?"
  },
  {
    "start": 1230.12,
    "end": 1238.56,
    "text": "Yeah I think especially in Germany right where a lot of companies don't have amazingly good internet connections though."
  },
  {
    "start": 1239.28,
    "end": 1242.62,
    "text": "so having something that requires cloud usage there maybe not."
  },
  {
    "start": 1245.46,
    "end": 1249.14,
    "text": "It's also not something that is completely out of scope because humans."
  },
  {
    "start": 1249.26,
    "end": 1259.9,
    "text": "Also have two hundred millisecond latencies for making decisions and only if you train really, really hard You can get it down to one hundred fifty or maybe a hundred milliseconds."
  },
  {
    "start": 1260.06,
    "end": 1265.76,
    "text": "these are the baseball people who have to catch balls And stuff like That."
  },
  {
    "start": 1265.86,
    "end": 1269.56,
    "text": "so So there Is A Like We See on The Human Baseline."
  },
  {
    "start": 1270.04,
    "end": 1270.12,
    "text": "What?"
  },
  {
    "start": 1270.24,
    "end": 1271.0,
    "text": "Actually Necessary."
  },
  {
    "start": 1271.3,
    "end": 1278.72,
    "text": "I guess Its Nice Way To segue Kind Of Into one of the other big topics in twenty-twenty six which is video models, world models."
  },
  {
    "start": 1279.54,
    "end": 1283.58,
    "text": "because I mean yes humans only react at like two hundred millisecond pace."
  },
  {
    "start": 1283.64,
    "end": 1293.26,
    "text": "right but we have a really good understanding of how the physical word around us works what certain actions has as i have...like..what the effect is of certain actions."
  },
  {
    "start": 1293.66,
    "end": 1305.46,
    "text": "Which basically means we can somehow internally predict What a certain action will do right and then when we're acting, were only really collecting correcting for this plan but."
  },
  {
    "start": 1305.92,
    "end": 1307.78,
    "text": "We have this inherent knowledge."
  },
  {
    "start": 1307.86,
    "end": 1315.26,
    "text": "I mean every new task Every new object that we are handling ,we have an intuition of okay how heavy is the subject going to be?"
  },
  {
    "start": 1315.38,
    "end": 1318.04,
    "text": "How it's gonna feel ?How its gonna behave Right?"
  },
  {
    "start": 1318.28,
    "end": 1324.98,
    "text": "And these sort things Is something That we basically now starting To get into When it comes to robotics."
  },
  {
    "start": 1325.46,
    "end": 1333.14,
    "text": "Okay so thats yet another piece context that humans do all the time, when I'm approaching a traffic light and it is yellow."
  },
  {
    "start": 1333.96,
    "end": 1339.66,
    "text": "I know that might need to stop five seconds before because they kind of anticipate what comes next."
  },
  {
    "start": 1340.42,
    "end": 1344.12,
    "text": "so this all goes under this label off world models right?"
  },
  {
    "start": 1344.54,
    "end": 1354.48,
    "text": "But like the term sounds very big seems like we're modeling everything in the world but obviously also only model are small part here."
  },
  {
    "start": 1354.64,
    "end": 1354.76,
    "text": "yeah"
  },
  {
    "start": 1355.28,
    "end": 1360.54,
    "text": "i think I mean, maybe we should start and like give a bit of an overview of what."
  },
  {
    "start": 1361.36,
    "end": 1363.34,
    "text": "What do you mean when we talk about world models?"
  },
  {
    "start": 1363.92,
    "end": 1388.14,
    "text": "And this can go... There's a wide variety but usually these days people mean is some type of model that basically based on the current state may be based in idea or what it supposed to happen look forward into future and predict certain thing of like how is this system going to evolve."
  },
  {
    "start": 1388.3,
    "end": 1392.68,
    "text": "And I think the most easy way to visualize and it's also the most common way."
  },
  {
    "start": 1392.84,
    "end": 1418.58,
    "text": "what we see these days, are video generation models where you essentially put in a language prompt so that you describe Imagine this scene and kind of give you a video what the scene looks like, how it evolves."
  },
  {
    "start": 1420.06,
    "end": 1443.88,
    "text": "And I think one of really big arguments for why roboticists are now interested in this is actually there's multiple arguments but i'll start with one which a video that has high fidelity and looks like this could actually be real."
  },
  {
    "start": 1444.92,
    "end": 1461.28,
    "text": "These models need to understand physics, in order to predict the very complex scene of how objects interact with each other so certain things fall down when they have no more support or pick up something."
  },
  {
    "start": 1461.72,
    "end": 1468.04,
    "text": "you basically to have two objects in close contact, right?"
  },
  {
    "start": 1468.44,
    "end": 1476.8,
    "text": "In order to efficiently model and predict these sort of things your video generation model basically needs some understanding or physics."
  },
  {
    "start": 1477.5,
    "end": 1488.84,
    "text": "And compared to vision language action models which basically have the vision-language backbone so what we now at ISF with JetGPT or Claude are something like that ,right ?"
  },
  {
    "start": 1489.26,
    "end": 1497.16,
    "text": "These models this vision language models they basically know about physics from from reading books and seeing images, right?"
  },
  {
    "start": 1497.66,
    "end": 1501.04,
    "text": "But a video model in order to really predict this."
  },
  {
    "start": 1501.2,
    "end": 1506.08,
    "text": "it needs more basic understanding of what physics is."
  },
  {
    "start": 1506.18,
    "end": 1509.02,
    "text": "How physics works how gravity works things like that."
  },
  {
    "start": 1509.72,
    "end": 1529.3,
    "text": "I think one of the big arguments for roboticists better properties and better representation of how a scene or the world looks like, because it basically has to learn that from data before."
  },
  {
    "start": 1530.32,
    "end": 1532.72,
    "text": "Okay so in essence what is a world model?"
  },
  {
    "start": 1532.76,
    "end": 1535.14,
    "text": "A world-model is something where I can predict future."
  },
  {
    "start": 1536.6,
    "end": 1539.24,
    "text": "The argument is humans do this all the time."
  },
  {
    "start": 1540.42,
    "end": 1543.32,
    "text": "even we needed for language models as well."
  },
  {
    "start": 1543.4,
    "end": 1546.54,
    "text": "right if Language models only just predict the next token."
  },
  {
    "start": 1547.54,
    "end": 1565.6,
    "text": "And so if I say, I'm going west starting from Munich to the city off and now I need to predict a city everybody would say well it's Augsburg because that is how this relation gets baked in until language model."
  },
  {
    "start": 1565.68,
    "end": 1566.24,
    "text": "as an argument"
  },
  {
    "start": 1571.7,
    "end": 1575.14,
    "text": "Roughly speaking, you could basically argue that the language model is a word model."
  },
  {
    "start": 1575.68,
    "end": 1578.1,
    "text": "But I think in terms of usefulness right?"
  },
  {
    "start": 1578.42,
    "end": 1590.82,
    "text": "i would argue That video model which Basically has A better understanding Of physics and how objects interact Because it needs that In order to predict what happens when two objects collide."
  },
  {
    "start": 1590.86,
    "end": 1598.02,
    "text": "Right there still It's usually or more useful world model."
  },
  {
    "start": 1598.24,
    "end": 1599.2,
    "text": "Let's call it that, yeah?"
  },
  {
    "start": 1599.48,
    "end": 1619.84,
    "text": "It is also a little bit over the top right because video model needs to predict pixels millions of values for every step but its amazing progress we have had in past few years while you can generate whole scenes and even longer sequences there really with high fidelity."
  },
  {
    "start": 1620.26,
    "end": 1622.9,
    "text": "But do we need all this?"
  },
  {
    "start": 1623.0,
    "end": 1623.58,
    "text": "Its good point."
  },
  {
    "start": 1623.68,
    "end": 1628.68,
    "text": "I mean theres kind two competing architectures i would say right."
  },
  {
    "start": 1628.86,
    "end": 1656.88,
    "text": "um one is the yeah plain video generation models where usually the actual generations still happens in a compressed latent space so you're not predicting full pixels or, But still, I mean the end is then basically you have a decoder which takes this compressed space and basically decodes that into pixels."
  },
  {
    "start": 1657.02,
    "end": 1659.64,
    "text": "And to something You can see... Which i mean."
  },
  {
    "start": 1659.7,
    "end": 1666.24,
    "text": "The big argument for That Is it's just super easy To debug because you Can look at the video?"
  },
  {
    "start": 1667.3,
    "end": 1669.9,
    "text": "If It looks good well Then your model kind of Works!"
  },
  {
    "start": 1670.22,
    "end": 1675.98,
    "text": "I guess the other competing architecture Is like inspired by Jepa."
  },
  {
    "start": 1676.66,
    "end": 1677.26,
    "text": "Where the idea?"
  },
  {
    "start": 1678.48,
    "end": 1686.02,
    "text": "Actually I don't have to learn this in the pixel space, but i can instead directly learn it in a compressed space."
  },
  {
    "start": 1686.5,
    "end": 1702.5,
    "text": "So instead of learning how to uncompress basically my compressed space into pixels... ...I just learned away from how to compress pixels into very good compressed space and then really only learnt by world model in the compressed space!"
  },
  {
    "start": 1706.3,
    "end": 1708.56,
    "text": "It's still the race of, and that is still open."
  },
  {
    "start": 1708.64,
    "end": 1709.78,
    "text": "Like what is better?"
  },
  {
    "start": 1710.32,
    "end": 1714.74,
    "text": "I think there are good arguments for both."
  },
  {
    "start": 1717.86,
    "end": 1738.86,
    "text": "one big thing is for sure how if you do it the JAPA way You can basically only use this as like a world model in the sense But it doesn't really give you useful output whereas If you have away off getting back to pixels you have a video generation model, which has basically a bunch of applications nonetheless."
  },
  {
    "start": 1741.18,
    "end": 1746.8,
    "text": "So just to sum it up so we have the videos themselves."
  },
  {
    "start": 1747.06,
    "end": 1750.06,
    "text": "they contain tonne information that don't necessarily need."
  },
  {
    "start": 1751.26,
    "end": 1757.82,
    "text": "things like shading and spatial thing from background not really interested in."
  },
  {
    "start": 1757.98,
    "end": 1761.34,
    "text": "We talked about state space before this all becomes now State."
  },
  {
    "start": 1765.78,
    "end": 1770.78,
    "text": "The idea is that you can have a lossy compression, right?"
  },
  {
    "start": 1771.08,
    "end": 1771.84,
    "text": "I think that's the point."
  },
  {
    "start": 1772.5,
    "end": 1779.3,
    "text": "You only keep certain learned concepts which are relevant for the task."
  },
  {
    "start": 1780.22,
    "end": 1785.04,
    "text": "but how do actually decide what to keep and how to compress?"
  },
  {
    "start": 1787.36,
    "end": 1788.42,
    "text": "Honestly through learning."
  },
  {
    "start": 1789.14,
    "end": 1793.82,
    "text": "so basically learn architecture such as the compression mechanism."
  },
  {
    "start": 1794.02,
    "end": 1795.26,
    "text": "it just very efficient."
  },
  {
    "start": 1796.62,
    "end": 1814.54,
    "text": "But I mean, i would argue that even the pixel video models right they have a certain way of compressing informations like shading or background texture because there it's also usually when you look at these videos."
  },
  {
    "start": 1814.62,
    "end": 1826.86,
    "text": "They have very high level detail in foreground and then There can be blurry background where you see okay model has basically some compressed version of what the background should look like, and then you basically just predicted it."
  },
  {
    "start": 1828.64,
    "end": 1828.86,
    "text": "Okay?"
  },
  {
    "start": 1830.08,
    "end": 1836.24,
    "text": "And so now let's say a rollout over to next second... We're not talking about minutes right?"
  },
  {
    "start": 1836.28,
    "end": 1840.14,
    "text": "we are talking about something that is in very near future."
  },
  {
    "start": 1841.68,
    "end": 1842.14,
    "text": "What Next?"
  },
  {
    "start": 1842.88,
    "end": 1848.38,
    "text": "Imagine this late vector or this compressed vector Or set of hundreds thousands pixels."
  },
  {
    "start": 1849.38,
    "end": 1853.14,
    "text": "Yeah, I think that's probably a good point to basically talk now about."
  },
  {
    "start": 1853.82,
    "end": 1855.9,
    "text": "how can we use these models in robotics?"
  },
  {
    "start": 1856.8,
    "end": 1862.24,
    "text": "And there is kind of roughly three big ways."
  },
  {
    "start": 1863.08,
    "end": 1881.82,
    "text": "One is artificial data generation so... We know how costly it is to collect tele-operation or human echo data And you can use these type of word models basically as a simulator and then learn in them."
  },
  {
    "start": 1882.16,
    "end": 1887.52,
    "text": "That's one option, um...and the other two I guess bigger options are."
  },
  {
    "start": 1888.26,
    "end": 1900.2,
    "text": "if you now know how the world is supposed to evolve over the next seconds You can basically use that knowledge and then learned what we call an inverse dynamics model where the idea really Okay."
  },
  {
    "start": 1900.3,
    "end": 1902.84,
    "text": "Given that this is supposed to be the future trajectory, right?"
  },
  {
    "start": 1902.96,
    "end": 1905.84,
    "text": "What do I have to do now in order to make it happen?"
  },
  {
    "start": 1906.56,
    "end": 1917.54,
    "text": "and one big argument is essentially learning such an inverse dynamics model when you have a video model or if your model predicts its future very well."
  },
  {
    "start": 1917.92,
    "end": 1929.42,
    "text": "It's much easier task than basically being at current point of view and deciding okay what should i need given that i know what i want but not give me what my future will look like."
  },
  {
    "start": 1930.72,
    "end": 1934.1,
    "text": "So that means the future becomes an input, right?"
  },
  {
    "start": 1934.14,
    "end": 1936.18,
    "text": "It becomes part of the input space."
  },
  {
    "start": 1936.3,
    "end": 1936.66,
    "text": "Correct"
  },
  {
    "start": 1936.7,
    "end": 1937.92,
    "text": "Yeah"
  },
  {
    "start": 1937.96,
    "end": 1940.58,
    "text": "Right so and we do this also every day."
  },
  {
    "start": 1940.78,
    "end": 1951.7,
    "text": "maybe We have An image in our mind how The kitchen should look like after we clean up And then suddenly Then the task is to align reality with what we had in mind."
  },
  {
    "start": 1951.78,
    "end": 1955.44,
    "text": "yeah, and then there's these people who cannot really."
  },
  {
    "start": 1957.18,
    "end": 1960.5,
    "text": "They don't haven like inner pictures or in an imagination."
  },
  {
    "start": 1961.24,
    "end": 1973.32,
    "text": "And maybe that's how to draw the parallel, they can still do the same things and somehow don't materialize it in images but have this vague notion of what should be then the case?"
  },
  {
    "start": 1973.48,
    "end": 1978.66,
    "text": "That is maybe one could think about these latent space models."
  },
  {
    "start": 1979.36,
    "end": 1992.38,
    "text": "I think its kind important because most video action models or world action malls that we see these days, they actually don't necessarily go all the way to predicting pixels again."
  },
  {
    "start": 1992.6,
    "end": 2005.68,
    "text": "So usually you keep doing the prediction of what is supposed to happen but you keep it in a compressed space and then basically use this knowledge with another model to get your actions."
  },
  {
    "start": 2006.2,
    "end": 2013.58,
    "text": "It's like from computational point-of view Predicting pixels entirely is very costly."
  },
  {
    "start": 2013.74,
    "end": 2024.44,
    "text": "It's like requires beefy big GPUs and if you can skip that, That's basically an advantage because he just have a faster execution frequency which I mean we talked about before."
  },
  {
    "start": 2024.58,
    "end": 2025.66,
    "text": "this important for robotics."
  },
  {
    "start": 2026.46,
    "end": 2028.14,
    "text": "Yeah And You touch upon it Before."
  },
  {
    "start": 2028.34,
    "end": 2030.44,
    "text": "let's Let's Just Take A Bit."
  },
  {
    "start": 2031.62,
    "end": 2032.36,
    "text": "This The Thing."
  },
  {
    "start": 2032.42,
    "end": 2040.02,
    "text": "You Know What Will Happen in the Next Six to Twelve Months There Because One might be the better architecture."
  },
  {
    "start": 2040.54,
    "end": 2047.8,
    "text": "The other one might be where we have already a business model around it, so in the end what will win there?"
  },
  {
    "start": 2048.92,
    "end": 2054.28,
    "text": "Honestly I think that's like looking into the crystal ball...I wouldn't dare to say no!"
  },
  {
    "start": 2056.1,
    "end": 2061.639,
    "text": "Both camps have good arguments and followers who are very convinced."
  },
  {
    "start": 2061.76,
    "end": 2063.08,
    "text": "this is how you should do it."
  },
  {
    "start": 2063.94,
    "end": 2073.239,
    "text": "these days we're seeing more models that basically go to pixels eventually, simply because there is a good commercial argument for it."
  },
  {
    "start": 2074.46,
    "end": 2078.219,
    "text": "But I wouldn't dare say one camp will definitely win in six months."
  },
  {
    "start": 2079.8,
    "end": 2091.62,
    "text": "One argument could be... There's tons of money and just video generation so maybe they are in the position then also tackle robotics alright?"
  },
  {
    "start": 2094.159,
    "end": 2098.5,
    "text": "So Okay, so we talked about images."
  },
  {
    "start": 2099.5,
    "end": 2109.2,
    "text": "We talked about memory and we talked About using future predictions as part of a memory maybe one thing around that."
  },
  {
    "start": 2110.1,
    "end": 2112.64,
    "text": "obviously there's not only One Future right?"
  },
  {
    "start": 2112.7,
    "end": 2117.06,
    "text": "So There is the this old example Of people who have worked on This ten years ago."
  },
  {
    "start": 2117.14,
    "end": 2126.62,
    "text": "I said okay you put A pen On its tip it's impossible to predict where it falls, and then people were talking about mode collapse."
  },
  {
    "start": 2126.7,
    "end": 2128.94,
    "text": "And blurry images and so on?"
  },
  {
    "start": 2129.02,
    "end": 2132.42,
    "text": "So how is that actually done with image models?"
  },
  {
    "start": 2132.72,
    "end": 2134.42,
    "text": "Are there any restrictions?"
  },
  {
    "start": 2136.12,
    "end": 2151.66,
    "text": "I think this kind of like the third way or for you could use these models in robotics right Where You basically look at a bunch of potential futures which one you want and basically use the actions according to that."
  },
  {
    "start": 2151.74,
    "end": 2165.54,
    "text": "So, more used model as a simulator planner where you can weigh options of what is actually future I want then act accordingly?"
  },
  {
    "start": 2166.56,
    "end": 2168.86,
    "text": "But it's all still very much sampling based."
  },
  {
    "start": 2169.9,
    "end": 2174.62,
    "text": "so somehow make decision on how things roll out."
  },
  {
    "start": 2175.08,
    "end": 2181.2,
    "text": "a wave, you're standing at the ocean and there's a wave coming to you."
  },
  {
    "start": 2182.2,
    "end": 2188.8,
    "text": "The grand scheme of things that is probably quite linear roll out but then the devil in details how it actually looks like?"
  },
  {
    "start": 2189.14,
    "end": 2189.3,
    "text": "Yeah"
  },
  {
    "start": 2191.4,
    "end": 2191.66,
    "text": "I think."
  },
  {
    "start": 2191.9,
    "end": 2193.98,
    "text": "so this basically comes back."
  },
  {
    "start": 2195.22,
    "end": 2205.56,
    "text": "what modern architecture of these models are which some type of diffusion or flow matching model where even if we have very multi-modal distribution."
  },
  {
    "start": 2206.54,
    "end": 2215.34,
    "text": "I mean, what these models are very good at and that's kind of a difference to what we had before which if you like now model this as some type of Gaussian right?"
  },
  {
    "start": 2215.42,
    "end": 2216.98,
    "text": "You get basically what you just talked about."
  },
  {
    "start": 2217.04,
    "end": 2224.7,
    "text": "Just the blurry image where wherever you have these multiple possible futures... ...you just get some sort of average of them."
  },
  {
    "start": 2225.6,
    "end": 2230.22,
    "text": "but i mean with these models in what they're like flow matching paradigm is really good."
  },
  {
    "start": 2230.28,
    "end": 2235.22,
    "text": "it is basically add these decision points decide for one of them follow that."
  },
  {
    "start": 2235.66,
    "end": 2245.02,
    "text": "And it might be that if you sample two or three times, so basically get different versions of it but at least you got like one consistent version whenever your sample which is a big plus"
  },
  {
    "start": 2246.0,
    "end": 2257.62,
    "text": "and also the reason why something like Will Smith eating spaghetti produces in video that actually shows something although there's an infinite number of versions of Will Smith out there!"
  },
  {
    "start": 2257.88,
    "end": 2258.0,
    "text": "All"
  },
  {
    "start": 2259.26,
    "end": 2259.44,
    "text": "right."
  },
  {
    "start": 2259.76,
    "end": 2261.5,
    "text": "So Video as One Thing."
  },
  {
    "start": 2262.76,
    "end": 2263.64,
    "text": "What Else Do We Have?"
  },
  {
    "start": 2264.36,
    "end": 2269.54,
    "text": "There's obviously the whole question of, what do you need to actually run robotics?"
  },
  {
    "start": 2269.66,
    "end": 2271.9,
    "text": "And that's probably far more than vision."
  },
  {
    "start": 2272.5,
    "end": 2272.96,
    "text": "Yeah for sure."
  },
  {
    "start": 2273.08,
    "end": 2278.48,
    "text": "I mean we can maybe talk a little bit about other types of modalities."
  },
  {
    "start": 2279.36,
    "end": 2290.84,
    "text": "We really see how people are experimenting with like with more inputs then just The robot proprioception which is joint readings and images videos."
  },
  {
    "start": 2290.92,
    "end": 2298.6,
    "text": "basically We started seeing people adding depth, for example to these models or a sense of touch."
  },
  {
    "start": 2299.16,
    "end": 2299.68,
    "text": "Or force."
  },
  {
    "start": 2299.86,
    "end": 2304.96,
    "text": "so and I mean especially when we think about interacting with the environment right?"
  },
  {
    "start": 2305.2,
    "end": 2324.68,
    "text": "And like moving objects are handling certain soft objects like cables or rope something that you could actually argue this senses is much more important as human because There's a lot you as human can even do blindly, right?"
  },
  {
    "start": 2324.9,
    "end": 2326.02,
    "text": "Without seeing it."
  },
  {
    "start": 2326.14,
    "end": 2338.86,
    "text": "I mean if your in front of the door and have to open it with your key You just through your sense of touch basically feel where the keyhole is And without any sense of seeing."
  },
  {
    "start": 2338.96,
    "end": 2351.3,
    "text": "So arguably these senses are equally important or almost as important And we see more and more models basically going into that direction, adding additional modalities."
  },
  {
    "start": 2351.52,
    "end": 2356.64,
    "text": "Basically seeing if we can boost the performance of these models by adding such additional modality."
  },
  {
    "start": 2358.98,
    "end": 2363.82,
    "text": "So I mean it seems like humans have some kind of redundancy there."
  },
  {
    "start": 2363.94,
    "end": 2365.9,
    "text": "so you always learn with all the senses."
  },
  {
    "start": 2365.98,
    "end": 2368.92,
    "text": "now switch off one sense and still do it?"
  },
  {
    "start": 2369.28,
    "end": 2369.38,
    "text": "Yeah!"
  },
  {
    "start": 2371.12,
    "end": 2375.14,
    "text": "And i think probably would like to chime in here."
  },
  {
    "start": 2375.8,
    "end": 2382.34,
    "text": "This is also an extremely, extremely interesting problem from the mechanical engineering point of view."
  },
  {
    "start": 2382.9,
    "end": 2383.68,
    "text": "Because in the end."
  },
  {
    "start": 2383.76,
    "end": 2385.54,
    "text": "now we're getting to the next bottleneck."
  },
  {
    "start": 2385.62,
    "end": 2397.48,
    "text": "how do you actually create robots that are robust enough to withstand production environments weeks on and sensors not breaking or drifting?"
  },
  {
    "start": 2398.8,
    "end": 2402.02,
    "text": "so we have a lot need for mechanical engineers."
  },
  {
    "start": 2405.26,
    "end": 2415.62,
    "text": "Everybody is encouraged to not only just switch to AI modeling, but also really solve these completely unsolved problems."
  },
  {
    "start": 2416.24,
    "end": 2422.9,
    "text": "Really big need in the market and not enough products out there And obviously we're trying do our part on this."
  },
  {
    "start": 2424.54,
    "end": 2442.38,
    "text": "So looking back at this year you've been now very central running robot intelligence from a like technical point of view, making sure that we're doing the right things at the right time."
  },
  {
    "start": 2442.48,
    "end": 2445.06,
    "text": "So what did you see over the course of year?"
  },
  {
    "start": 2445.14,
    "end": 2446.68,
    "text": "What has changed in space?"
  },
  {
    "start": 2447.34,
    "end": 2459.2,
    "text": "I think one really important observation is that um The models keep getting better and better And-and i mean We have to be really careful about Um what battles we pick?"
  },
  {
    "start": 2463.66,
    "end": 2471.04,
    "text": "The model that we basically used at the beginning of the year Was completely dominated by the model."
  },
  {
    "start": 2471.08,
    "end": 2471.86,
    "text": "That were using now."
  },
  {
    "start": 2472.0,
    "end": 2482.42,
    "text": "so Every three months, we're getting new models with new capabilities through partners or Through just the open-source world."
  },
  {
    "start": 2482.54,
    "end": 2489.6,
    "text": "there's a lot off There's a load of things happening there and And we see that this space is moving really fast."
  },
  {
    "start": 2489.7,
    "end": 2495.34,
    "text": "So it's very amazing how how quick the progress is these days."
  },
  {
    "start": 2496.74,
    "end": 2499.12,
    "text": "And I mean, in terms of Robco right?"
  },
  {
    "start": 2499.46,
    "end": 2505.58,
    "text": "I think our goal really to get this models to a state where we can actually deploy them."
  },
  {
    "start": 2505.86,
    "end": 2512.5,
    "text": "twenty-four seven in an operation and as you said hardware's basically just that important ssd intelligence around it."
  },
  {
    "start": 2514.28,
    "end": 2519.08,
    "text": "Yeah there are bunches interesting problems were working on solving"
  },
  {
    "start": 2521.84,
    "end": 2522.0,
    "text": "Right."
  },
  {
    "start": 2522.06,
    "end": 2541.5,
    "text": "And maybe one should say that this is not really an ivory tower kind of operation, it's customers specific things out on the factory floor have a real need and we're banging our heads against them because they've made a lot progress."
  },
  {
    "start": 2541.94,
    "end": 2548.0,
    "text": "so tell me what do you think about today?"
  },
  {
    "start": 2549.66,
    "end": 2554.86,
    "text": "How do you learn as a group and how do you structure this work?"
  },
  {
    "start": 2555.36,
    "end": 2561.74,
    "text": "So, I mean we basically have pretty rigorous framework for experimentation where."
  },
  {
    "start": 2562.78,
    "end": 2567.96,
    "text": "We have a continuous radar going on of okay what is the rest of the world doing?"
  },
  {
    "start": 2568.08,
    "end": 2583.28,
    "text": "because i mean we have to basically keep up-to-date with other people are doing so were like going through research papers what we can learn from them, and then in terms of how we progress ourselves."
  },
  {
    "start": 2583.54,
    "end": 2586.78,
    "text": "We basically come together as a group two times a week."
  },
  {
    "start": 2587.24,
    "end": 2600.64,
    "text": "Uh...we propose experiments um ...and then we execute them ,we track them .We analyze them And over the course time through these discovery loops we get better."
  },
  {
    "start": 2604.1,
    "end": 2612.46,
    "text": "In a year from now, we can basically go back with an AI agent actually and just like sort through okay what experiments have we done?"
  },
  {
    "start": 2612.68,
    "end": 2613.92,
    "text": "Have we tried this before."
  },
  {
    "start": 2614.34,
    "end": 2615.26,
    "text": "What have we learned here?"
  },
  {
    "start": 2616.54,
    "end": 2623.34,
    "text": "so making this type of structured experimentation part of the daily work routine is was really key at Robco."
  },
  {
    "start": 2624.26,
    "end": 2627.84,
    "text": "So really applying the scientific method as part of an engineering org"
  },
  {
    "start": 2628.0,
    "end": 2630.06,
    "text": "yeah I think you have to do that."
  },
  {
    "start": 2634.2,
    "end": 2636.32,
    "text": "put your hypothesis to the test, right?"
  },
  {
    "start": 2636.96,
    "end": 2643.04,
    "text": "You come up with certain ideas or you see something in another paper and then you say okay let's try this maybe."
  },
  {
    "start": 2643.52,
    "end": 2644.12,
    "text": "This helps us."
  },
  {
    "start": 2644.68,
    "end": 2647.46,
    "text": "uh...you have a hypothesis- you do an experiment around it."
  },
  {
    "start": 2648.06,
    "end": 2650.62,
    "text": "Uh..You confirm that you deny it And basically move on."
  },
  {
    "start": 2650.66,
    "end": 2656.2,
    "text": "Then you gain knowledge through That and apply It To The Product We In The End Wanted Ship To The Customer."
  },
  {
    "start": 2657.76,
    "end": 2658.08,
    "text": "All Right."
  },
  {
    "start": 2658.88,
    "end": 2661.88,
    "text": "So all of this is context Of also new product development."
  },
  {
    "start": 2662.0,
    "end": 2662.8,
    "text": "Nothing Is Fixed."
  },
  {
    "start": 2663.98,
    "end": 2671.92,
    "text": "working end-to-end on the combined hardware and software solution, making lots of progress over the course this year."
  },
  {
    "start": 2672.88,
    "end": 2675.5,
    "text": "I guess next year will be at least as interesting."
  },
  {
    "start": 2675.54,
    "end": 2682.3,
    "text": "we're getting closer to actual production deployments that i hope can talk about here really soon."
  },
  {
    "start": 2683.48,
    "end": 2693.72,
    "text": "so with it's a good round up And how it ties in that, you know apparently doesn't always predict forty two."
  },
  {
    "start": 2694.02,
    "end": 2698.82,
    "text": "It predicts the future and it informs what the robot actually does."
  },
  {
    "start": 2699.36,
    "end": 2702.62,
    "text": "so yeah really exciting progress over the course of this year."
  },
  {
    "start": 2703.08,
    "end": 2705.16,
    "text": "think three years ago nobody talked about world models."
  },
  {
    "start": 2705.24,
    "end": 2709.06,
    "text": "yet now we're at a point where it says okay is This The Real Thing or Is There?"
  },
  {
    "start": 2709.32,
    "end": 2711.68,
    "text": "The Next Thing In Vicinity?"
  },
  {
    "start": 2713.22,
    "end": 2716.52,
    "text": "If It Happens You Know We Have It On The Horizon Definitely As Well."
  },
  {
    "start": 2717.04,
    "end": 2717.66,
    "text": "So Yeah With That."
  },
  {
    "start": 2718.28,
    "end": 2721.02,
    "text": "Thank you, Felix for joining the podcast."
  },
  {
    "start": 2721.88,
    "end": 2722.26,
    "text": "Yeah"
  },
  {
    "start": 2722.82,
    "end": 2724.84,
    "text": "and yeah hope to see your next time."
  },
  {
    "start": 2725.6,
    "end": 2728.24,
    "text": "always click on like subscribe."
  },
  {
    "start": 2729.02,
    "end": 2731.74,
    "text": "come back when it says Rob talk."
  },
  {
    "start": 2733.42,
    "end": 2736.18,
    "text": "this is Rob Talk a podcast by Robco."
  },
  {
    "start": 2737.02,
    "end": 2738.24,
    "text": "Subscribe on Spotify"
  },
  {
    "start": 2738.46,
    "end": 2739.62,
    "text": "Apple Podcasts and"
  },
  {
    "start": 2739.7,
    "end": 2740.06,
    "text": "YouTube."
  },
  {
    "start": 2740.76,
    "end": 2742.58,
    "text": "See You On The Factory Floor."
  }
]