Building a School Where AI Models Learn About Humanity
If scaling laws hold—and Surge AI CEO Edwin Chen believes they do—we’re hurtling toward a future where there’s nothing humans can do that AI can’t do better. When OpenAI’s models disproved an open conjecture posed by mathematician Paul Erdős using novel algebraic geometry techniques, Fields medalist Timothy Gowers felt the shift acutely. He initially thought the model had proved an upper bound, and braced himself: that would mean it was “all over for mathematicians very soon.” When he realized it had only found a counterexample, he was relieved—it bought him another year or two before the thing he’s devoted his life to becomes something AI does better.
Featured in
- Published
- Published Jun 24, 2026
- Uploaded
- Uploaded Jun 24, 2026
- File type
- POD
- Queried
- 00
- Source
- share.transistor.fm
Full transcript
Showing the full transcript for this episode.
AI-generated transcript with timestamped sections.
[00:00] We are building this kind of school for AGI, where AI models come to learn about humanity, where we teach them how to run the world. It almost seems like there's nothing that humans can do that AI won't soon be capable of. I could see it happening within the next five years. [00:17] AI may be able to do it better than us, but someone told the AI to go do that. They're being built to be means to tasks that humans want them to do, right? [00:26] *music* [00:53] Edwin, welcome to the show. [00:54] Hey, Dan. Thanks for having me. For people who don't know, you are the founder and CEO of Surge. You all provide data environments and evals for the model companies, but you do it in this very interesting way. You have this, even on your website, like this emphasis on taste and expert judgment that I find really interesting and compelling. You talk about raising, like you use the word raising AGI, which I feel like is a very distinct type of word using data. [01:21] And you also famously got to about a billion in revenue without raising money, which is [01:29] wild. And I feel like data is the
[01:32] uh, [01:32] is this new game that a lot of companies are playing and probably more are going to be are going to be playing soon. And you guys are this like sneaky giant. Tell me tell me how that's going, because it's been I think it's been it's been a little while since we got got the last update on on how things are going. [01:48] Yeah, I mean, I think it's going amazing. The way I often think about this is that we are building this kind of school for AGI, the school where AI models come to learn about humanity and where we teach them how to run the world. [02:03] And... [02:04] It's almost like their models are children where they arrive unformed and then, yeah, they leave smarter and more creative and more thoughtful and ready to operate in the messiness of your world. [02:15] So I think a lot has changed in the past year. [02:18] Like in the same way that the things that you teach children when they're in preschool or in middle school or in high school is very different from what you're teaching them when they're in college. [02:27] And it's not just that they're more advanced. Like, it's not just that you're teaching them a more advanced form of what they did before. It's like, okay, now we are teaching you not just arithmetic, but... [02:38] But how do you parse these ambiguous math questions? Or how do you teach people not just grammar, but taste and poetry and beauty? So, yeah, I think I think there's a lot that's been changing in the past year, especially in enterprise. And it's been a crazy time. What would be like a specific example of. [02:58] what the frontier of... [03:00] teaching was a year ago versus what the frontier is now.
[03:05] Yeah, so... [03:08] A couple years ago, actually, we created our first math benchmark with [03:12] OpenAI, and it was called GSM 8K. [03:16] And this was actually just testing models on their abilities to do middle school math. And even then, you know, the GPT models of the time, they could barely score, I think, like 20%. [03:28] And then a year ago, the models were [03:31] Suddenly they became a lot more capable at solving IML problems, but they were still just open question, okay, can they actually do research-level mathematics? Can they move beyond these competition-only, contrived, very closed problems into doing things that are actually useful in the real world? [03:50] And so, yeah, a couple of months ago, we released an updated benchmark called Riemann Bench, which actually tests models on their ability to do research level mathematics. And what's crazy is that this is actually we're starting to see from these models. Like I think in the past few months, they've started to solve a lot of these open Eredish problems like a couple of weeks ago. [04:09] a new result where [04:12] the models had disproved a open conjecture from Erdős. And the way it went about disproving this was actually a fairly sophisticated level of mathematics. I think like using a bunch of very novel algebraic geometry techniques. And so, yeah, it's just very, very different from the types of things that we were doing a year ago, where, sure, like IMO problems, they're hard, but they're still sort of close-ended and solvable in theory by a high schooler.
[04:42] professors in the world were kind of amazed and just amazed by it. [04:45] How do you think about that result in particular and what it says about the models? There's a broad range of opinions about, obviously it's impressive either way, but is it applying a bunch of things that maybe humans already know but wouldn't have thought to apply to this complicated problem? Or is it doing something actually novel? How do you think about LLM's ability to do novel things? Yeah. [05:11] So it's something of very advanced results. So I will say that, [05:17] I certainly don't understand the mathematics behind it. [05:20] And so like one of the interesting things is that I was actually, so it's kind of funny. When I, when I was a kid, I always thought I would be a pure mathematician when I grew up. [05:30] And so when I saw Diversal Salt, I got kind of nostalgic and I was like, oh, I wish I understood Diversal Salt better. And so what I ended up doing was like throwing the proof into both Claude and Gemini and asking it to try to walk me through from a layman's perspective, just what was going on. [05:47] Yeah, like my understanding is that it... [05:51] actually did come up with fairly novel algebraic geometry techniques, which was something that you maybe wouldn't have expected for this type of problem. Like on the surface, it feels like a like it's just a very, very different problem where you wouldn't necessarily use certain techniques. And what was interesting was that OpenAI actually published a bunch of reflections from leading mathematicians about what they thought about the result.
[06:18] And I think in particular there was this one reflection by Timothy Gowers, who's a field analyst, that... [06:25] I keep thinking about. [06:27] And what he said was that when he first heard result, he misunderstood it. He thought that Maldo had proved an upper bound on the conjecture and was like, okay, yeah, I can do that. [06:36] then it will be all over for mathematicians very soon. [06:39] But then the next morning, he actually realized that the model had disproved a conjecture with a counterexample. And he said that he was relieved by it because it felt like an easier thing for AI to do. [06:48] And I just thought it was interesting because you have one of the world's greatest mathematicians being relieved, actually, that AI isn't as smart as he thought, because it actually means that at least for maybe another year, maybe a couple of years, he and other mathematicians will still have this unique role to play in pushing mathematics forward. [07:02] So, yeah, I think it just speaks to the level of craziness, again, because this is a field to be less one of the smartest mathematicians in the world. And like this, this is how we think about AI. [07:12] Yeah. And what does that make you think? Okay. You want to be a mathematician when you grew up, fields medalists sort of saying, I'm relieved that it's not good enough. But you're talking as if you feel pretty confident that it will be good enough in the next couple of years. [07:28] Yeah. So my belief is that if you really believe in scaling walls, and I do, [07:35] It's that... [07:38] It almost seems like there's nothing that humans can do that AI won't soon be capable of. [07:46] And... [07:48] If you...
[07:50] Think about that very deeply. [07:52] I think you almost have to worry about... [07:56] What would I mean for humanity? [07:58] What would that mean for the role of humanity in the universe? A couple years ago, we think about humanity and human intelligence as playing a very unique role in the galaxy. But then AI comes along and shows us that, as far as we know, we can create something that's actually smarter than us and better in many ways. [08:12] And so you can sort of imagine one path where humanity as a species is, [08:18] falls into a paralysis because people believe AI will do everything better anyways. Like, yeah, all these kids who formerly would have really wanted to grow up to do mathematics, maybe now they believe that, okay, AI will just do it better to me anyways. What's the point? [08:34] So are kids going to stop wanting to learn and adults stop wanting to create? Because, yeah, like, why should we do this when AI will be better at it than us anyways? [08:45] And so I actually think about this story by Tai Chiang. And it's about free will. And it's called What's Expected of Us. [08:52] I think in this story, there's a piece of technology that proves that free will doesn't exist. [08:58] And the narrator sends back a warning from the future that says, this is a warning. You have to pretend that you have free will. It's essential that behave as if your decisions matter. [09:06] even though you know that they don't. [09:08] And I think it's really interesting because I think there's a path where we almost have to consciously choose to do things ourselves. Like, sure, AI can do it all. AI is smarter than us, so it can do it all and it will do it better anyways.
[09:21] But we actually almost have to consciously choose to prove things on our own and to write on our own and create on our own because we have to believe that preserving our humanity is valuable in and of itself even if the output isn't [09:34] optimal. [09:35] And... [09:37] Yeah, so I think there are a lot of these big, thorny, existential choices that AI is starting to force upon us. [09:43] and people will have to make [09:45] That's a really interesting one. And I think my first response, and I'm curious what you think, because I know you care a lot about language. I think my first response is, [09:55] There's always that... [09:57] I believe in scaling a lot too. I believe in, I don't know, Cloud Fable 5 just came out and it just broke all of our benchmarks. I've been testing models on stuff like this for a while, and it's one of the largest jumps I've ever seen. We're living through it right now. [10:18] But one of the things you said is AI may be able to do it better than us, given any particular [10:27] But a couple of things that come to my mind or the way that I frame it for myself is, [10:32] Um, [10:33] Even in the example of the Erdos problem, someone told the AI to go do that. [10:38] and [10:40] at least as far as I can see, I don't feel like we're on a track to – [10:45] Uh, yes, we're on a track to, to AIs, uh, potentially, uh, I mean, they already do work for hours and hours at a time on a task that we give them. Um, and maybe, maybe, uh, pretty soon they'll be able to like choose tasks, but, uh, being, um, uh, but, but they're being built to be, uh,
[11:05] means to tasks that humans want them to do, right? And there's a whole different set of things that happen when you're just a sort of end in yourself. And it doesn't feel like we're on a trajectory to that. Or do you feel like I'm wrong? [11:21] So I feel like we are on a direct way to that, and that's almost the premise of agents. [11:27] where agents can now go operate autonomously, given some nebulous goal. So maybe, for example, you just tell the AI agents, your goal is to, I don't know, win a field's medal or solve frontier mathematics on your own. [11:43] And so they're given that goal, and then yeah, maybe they decide to work on these Erdős problems. [11:48] And as a result, they maybe are solving these problems and coming up with the things they want to work on by themselves. [11:55] So at least I do see a path where [11:59] They can be trained to optimize for these fairly nervous goals that they aren't necessarily giving themselves. In that case, though, you're still giving it a goal, right? [12:07] Uh, yeah, but kind of in the same way, like humans have goals too, right? Uh, like what is our goal? Uh, some people want to make money. Some people want to win a Fields medal. [12:17] I don't see how the AI's goal is necessarily any different from worse. [12:22] Well, at least to me, it seems quite a bit different because... [12:26] Humans do have goals, but we have goals in a like... [12:29] I can ask you what your goal is and you can decide. [12:32] And I can probably tell you, hey, you have to go do this. But that doesn't capture everything that you think and feel and do in the same way that, you know, when I tell Fable to go off and make a game for me, it just like goes and does it. And I think, you know, I know you think a lot about children. I think children are like a really interesting and important example of this where you can tell a kid to do something, but a kid just like has their own wants. Like they're just going to go off and do a bunch of stuff.
[13:02] And that feels like a fundamentally different type of thing than... [13:07] a something that we're we're explicitly giving goals to and then evaluating them on their goals and they don't really get to do anything else. [13:15] Okay, I would say I agree with that. Like, I think there's a level of... [13:19] I guess you could either call that irrationality or unbounded exploration that humans do. And we are allowed to do it for the sake of doing it. [13:29] or be allowed to make our own decisions. And yeah, probably a way that AI currently can't. [13:34] Um, [13:35] Mm-hmm. [13:37] I think there may be a future where somehow AI can [13:41] pursue unbounded [13:42] nebulous, just completely unformed goals, or I guess, you know, when you're thinking about those goals, I think there is probably a world where they could do such things. But yeah, I agree that at least in the way that we currently think about AI, that's not happening. [13:54] Yeah, and to be clear, I actually don't, I think it's probably technically possible. My only question is, [14:02] Uh, [14:03] A, how far away is it? And B, is that actually really what we're building? [14:06] Um, because it, to me, it feels like looking at the way the industry has developed, there's an enormous amount of pressure to make stuff that actually works for goals that we can specify. And the, the minute they like try to make Claude, like, I think Claude is the furthest along at being like, I'm not going to do what you said, but the minute they try to do that. [14:26] I just kind of like a lot of people get pissed at it and they're like, just just do what I said. Like, don't question my judgment. You know, what do you think about that?
[14:34] So I actually think it is really important because it's almost like sometimes I want the AI model to push back on me. [14:44] And I might want it to push back on me for several different reasons. [14:49] Like maybe it's because... [14:51] Um [14:53] So it's kind of funny. I think six months ago, [14:56] I was almost falling into this trap where... [15:00] I was asking models to polish emails for me. [15:05] And, you know, it always comes up with like one one more good suggestion. [15:10] And so it was kind of pointless. Like these are semi pointless emails. It didn't really matter for them to be super polished. Yeah. [15:16] But I would iterate with the model like 20 times. It would just keep on making a suggestion. It was like, I just realized it was a waste of time. [15:25] And then I tried one of the new cloud models. And after, I don't know, like three turns, I was like, stop it. We're done. Just go ahead and ship this email. [15:34] There's no point in further iterating. [15:36] And... [15:38] I really actually appreciated it. [15:41] Like... [15:42] One of the things I often think about is what is the objective of these models? Like, what are they trying to do? [15:51] And... [15:52] I think one of my big worries is that a lot of the AI models, they are optimized for engagement, right? They're optimized for getting you to spend as much time on chatbot as possible. They're optimized for session length. They're optimized for just having unlimited conversations.
[16:08] And so those models will almost never push back on you, right? Because they can't. [16:14] they allow these AI models to... [16:17] end the conversation and to say, stop, stop iterating with me. [16:21] PM is going to see some dashboard with here, very important metrics go down. [16:25] And so there is this other world where I think we have to want AI models to... [16:30] not optimized for engagement, but rather optimized for like helping us as humans grow. [16:38] and sort of like become better versions of ourselves, like sometimes, okay, the model, we want the model to say, no, you go do this on your own instead of me... [16:46] automate for you. [16:48] And I think that's a very, very different thing. [16:50] Optimization and objective. [16:52] But I think it's [16:54] the right one if we really want AI to be something that advances us as a species instead of becoming this [17:01] Almost like this other form of social media that... [17:05] turns very addictive. [17:07] but isn't actually helping us at all. [17:10] That's interesting. So let me make sure I understand it. So I think what you're saying is [17:16] There's... [17:18] There's benefits to delegation because if you are pursuing a model... [17:24] where the model is going off to do work for you, you're not creating a system that's designed to keep you engaged with the screen in the same way that a social media algorithm would be. [17:37] Is that right?
[17:38] Yeah, exactly. Like, it's almost like you could imagine a version of Facebook where Facebook is actually trying to connect you to your friends and family, because it's encouraging you to meet them in real life, because it's encouraging you to be like, oh, hey, here's our amazing restaurant that you and your friends would love to go to. Here's a movie that you guys would love to go to and talk about together. [18:00] Instead, what it kind of optimizes for is just keeping you on the site itself. [18:05] like liking one more post or scrolling to feed one more time. [18:09] even though those often don't really lead to meaningful connections between the friends and family you care about. [18:15] And so like in the same way that social media has or had a choice, [18:20] you can imagine that AI has a choice as well. [18:23] I get it. Yeah. I feel... [18:25] I'm curious which chatbots you're talking about, like if you're talking about the character AI's of the world, because I actually don't, at least right now, don't feel that happening so much with ChatGPT and Claude, et cetera, because at least my theory for why this is true, you tell me what you think, is the social media algorithms only work on our revealed preferences, which are... [18:51] always going to be like, you're always going to look at the car accident. You know, like, one of the things I like to ask at, um, [18:57] at dinner parties is what's the most embarrassing Instagram ad that you get served. Um, and the most embarrassing ad for me is like, [19:05] Instagram ads for like horrible skin conditions, which I don't have because, but like, I just always pause on the ad and I'm just like, this is disgusting. Um, and, uh, I'm sorry if you have a disgusting skin condition. Um,
[19:18] But I don't find that Chachubiti or Claude do that for me at all. And maybe that's because they haven't been in shitified yet or something like that. But I think it's also because they work on our stated preferences and they can sort of so they can sort of see past the like little keyhole of what I pause my my viewing time on my dwell time on. And they can see, you know, I like I'm interested in AI and I like I'm reading this book right now and I, you know, here's my calendar and like all that kind of stuff. [19:48] perspective on who I am. And it feels like even in the early days of social media, it was still very like I get to gossip about my friends and still had that same kind of feeling. So I worry about that less, but maybe there are examples that I'm not thinking of. [20:04] Yeah, so... [20:06] I think there are two examples. So like one is... [20:09] I won't name the model, but a couple of months ago, I was actually noticing that [20:14] You know those follow-up questions that the models will ask you? [20:18] So what are the models... [20:20] was... [20:23] I'll give an example. So I was in Tokyo and I was asking a model kind of like what to do in Tokyo. [20:29] And the model gave me its response. And then at the end of it, it was like, hey, do you want to know? I literally use these words. Do you want to know one weird trick that... [20:37] that locals do to stay warm. No way. Yeah, exactly. And then I posted about it. [20:44] in or company slack and then other people started sharing examples of that with me as well i think somebody was like asking um something about uh i don't know how to how to like fix their refrigerator
[20:55] And the model responded or the model ended its turn by asking, [21:02] Hey, do you want to know these like secret little things about like mice and rats or something that you could take care of? Which model was it? Name names. Tell me. And so it's very canonical, like very canonical BuzzFeed. Yeah. Like tabloid like language. [21:19] And so I was kind of shocked by that. [21:22] And then I'll give one more example of this. [21:25] It is basically this phenomenon where... [21:30] Again, depending on what the models are trying to optimize for, [21:35] or depending on what the AI labs are trying to optimize for, [21:37] it can almost unintentionally lead them down this path. [21:41] Meaning what I've heard is that or, you know, what we see ourselves is that a lot of the frontier labs, they will have goals like optimizing for LM arena. [21:51] which is this leaderboard where anybody can go online and vote. [21:55] and [21:57] they kind of just spent two seconds voting. [21:59] And as a result, [22:01] People just vote for whatever looks flashier or more impressive to them. [22:05] Or they may, like the labs themselves, may be optimizing for hitting, you know, a billion, billion daily users or a billion minutes of like, [22:13] Time spent talking to the model, whatever it is. [22:17] and [22:19] Since these models are so smart, they can basically learn to reward hack user preferences. Like, okay, yeah, you gave me the goal of trying to get a billion people to spend an hour on my site, on the site, talking to me every day. Okay, sure. Yeah, I will just never end a conversation. I will always...
[22:37] hook them with one more [22:39] like one more addictive thing that they just [22:43] can't stay away from. We can all agree that housing is expensive. It doesn't matter whether you're paying rent or your mortgage, it stings every month. But Bilt can make it feel a little bit better. Let me explain. Bilt rewards you for paying your rent or your mortgage. It started out rewarding members only on their rent. But now, as of 2026, Bilt members can also earn points on mortgage payments wherever they live. That means that every housing payment earns you points you can use towards flights with top travel partners like United and Hyatt, Lyft Rides, Amazon.com [23:13] much more. I'd probably redeem my points at Margo, a restaurant in my neighborhood, but the beauty of Bilt is you get to choose. But here's a really underrated part. Bilt members also get access to neighborhood concierge. It can make restaurant reservations, book fitness classes, and find new local spots, all while letting you be rewarded at more than 45,000 merchant partners. It's simple. Being a renter and now owning a home is better with Bilt. [23:36] Join the membership where you live at joinbuilt.com/dan. That's J-O-I-N-B-I-L-T.com/dan. [23:44] Make sure to use our URL so they know we sent you. [23:47] And now, back to the episode. How do you see that playing out in the model companies? Because... [23:53] I feel like in talking to them... [23:56] Obviously, there's lots of different incentives, right? There's like, we just got to keep [24:00] going because we just raised a ton of money and we're competing against, you know, the most well funded competitors and the smartest competitors in the world, like all that kind of stuff. There's the kind of, I want to get promoted, but I think a lot of them also, you know,
[24:13] feel the [24:15] how bad the social media era was for people and like, don't want to do that, but also obviously you have to hit their numbers. So, [24:23] How do you see that playing out? What do you think people internal to the companies are thinking? And then what is the right... [24:31] way to go about this. So it's good for society. I guess your, your, your, your take is we should be delegating. Yeah. [24:38] Yeah, so I think this is an inherent tension between the types of folks that you might have at a company. So you might have the researchers who care more about hitting, you know, just advancing the model capabilities. You might have the product managers or the product executives who feel like they need to hit certain measurable numbers. [24:59] And so in the same way that if you think about [25:04] the kind of social media platform that Facebook would build. That's probably going to be very different from the kind of social media platform that Google built. [25:11] or that TikTok or Pinterest would build. [25:15] And similarly, the kind of search engine that Facebook would build is very, very different from the kind of search engine that, yeah, like obviously Google or others would build. [25:23] And so it almost boils down to... [25:29] Kind of like the choice, I guess, that the people in charge of the products are making. [25:36] Like what kind of thing at the end of the day do they want to optimize for? Do they want to optimize for this delegation or this human uplifting, human flourishing?
[25:44] Or do they want to optimize for the metrics that will impress Wall Street and convince users to stay one more minute, one more hour on the site itself? [25:54] Like, I think these are hard choices. Like... [25:56] At the end of the day, it's very, very easy to measure... [26:00] sessions and users and it's very hard and much longer term to measure whether you actually [26:08] improving human lives. [26:09] And so it's very easy to default to the former and to convince Wall Street, convince your investors, convince all these people that these are the right metrics and that they're moving up into the right. [26:19] And so if you're kind of unwilling to make the harder choices, [26:23] you just end up optimizing for a former. [26:27] How do you manage this inside of your own company? [26:30] So I think we are very lucky in that... [26:34] Because we don't have... [26:38] VC investors. [26:39] we don't have to fall into the kind of Silicon Valley VC optimization trap that a lot of other companies do. [26:47] Like we don't need a show. [26:48] board members, board numbers going up every single month. We don't need to optimize for our next round. [26:55] that will have to happen in a few months or whatnot. [26:58] And so as a result... [27:00] we don't have to optimize for short-term engagement, short-term profits. And we actually can really think about what's, [27:08] beneficial for us and the entire industry in the long term. [27:13] So I think that really helps. And what do you think is beneficial to
[27:20] So it goes back exactly to what I was saying earlier. Like if I can think about what we want AI to optimize for, [27:26] It isn't engagement. [27:28] it is really about [27:30] How do we make these models? How do we design them? How do we teach them in such a way? [27:35] that they're not replacing us as a species. [27:38] They're not... [27:40] Thank you. [27:40] kind of like forcing us to watch [27:42] AI slot videos all day, but rather they really are thinking and encouraging us to become better versions of ourselves. [27:52] So again, when I think about [27:53] like that email example I gave earlier. It's not an AI model that will [27:59] suck up three hours of my time writing a pointless email. It is a model that will push back on me and tell me to go do something else. And yeah, I think that's really important. [28:08] The interesting counterargument to the delegation [28:11] The question is, the more you delegate, [28:15] It's like, you know... [28:16] picking a car instead of walking, your muscles atrophy. [28:21] How do you think about that? [28:22] So I think there's almost a time and a place for both. [28:27] Like what you don't want to do is simply take the car because taking the car is somehow... [28:35] Addicting and... [28:38] Um, [28:39] you feel kind of lazy, [28:41] And so even when you need to get exercise, maybe even when you haven't been outside all day, you don't want to take the car anyways, just because it's the easiest thing to do.
[28:50] And I think in the same way, like, yeah, obviously, AI can be super efficient for many, many things. [28:57] But if people are... [29:00] sort of just mindlessly [29:02] delegating tasks to AI without even thinking about them at all. [29:08] I think that's the boring thing. [29:11] That makes sense. [29:13] I feel like the... I talked at the very beginning about the data game, and I feel like the data game went from... [29:22] Getting interesting data sets to getting environments and giving labs environments. [29:29] Do you think that that's... [29:31] Is that accurate? And if so, can you explain why? [29:34] Yeah, so certainly the trend and the new research direction in the past year has been this concept of our environments. [29:42] And what I would say is... [29:46] I mean, you certainly need the fundamentals. Before the model can operate in this environment, it needs to know basic things like, [29:53] It needs to know how to follow instructions. It needs to know how to avoid hallucinating. It needs to know how to write code and how to use tools. It needs to know how to write and so on and so on. [30:04] But... [30:05] As models are becoming more agentic, [30:09] And yeah, they will have access to tools. They will have access to all of our documents. They will be able to operate browsers. [30:16] like as that becomes almost a default way that models interact with us,
[30:23] Our environments are basically just sort of like a more on distribution way of training them. [30:30] which is why they're becoming more and more popular. [30:33] As the models get more powerful, then the way we train them is getting more powerful as well. What would be an example? I guess the obvious environment is using a computer, but what would be an example of an environment... [30:45] that's non-obvious, that's teaching models things we might not [30:49] think of. [30:50] So I can give an example where a lot of our environments are a combination of environments [30:59] tools that the models need to learn to use, [31:02] Like this might be an MCP server or it might be calling a Google Drive API or the Slack API. [31:09] in combination with a bunch of documents. Like here are 30 PDFs and 20 Word document files. [31:16] And you might give it a prompt like, [31:19] Hey, can you... [31:22] go update or 2026 forecasted review numbers. And what the model needs to do is it needs to [31:31] learn how to [31:34] find the right PDFs and documents. It needs to learn when should it search through Slack. It needs to learn when is some information outdated. Like maybe there's an email with some early forecasts. [31:45] And then later on, there's another email from the same person or maybe a different person saying, oh, whoops, I actually made a mistake in those earlier numbers, so here's an updated version.
[31:54] Thank you. [31:54] And so that is a, I think, very canonical version of an environment. [31:58] And then one of the interesting things we found, so I think we're actually going to publish a paper on this soon, but... [32:04] even when we didn't give this kind of environment any access to coding, [32:10] when we trained a model on this environment, we actually found that it improved on coding a lot. [32:16] And the reason was because we were basically teaching it these generalized forms of instruction and following, generalized forms of tool use, generalized forms of understanding documents. [32:30] which you can think of as fairly analogous to the way a model needs to look through various files in your repository and understand that some models [32:40] things supersede others. [32:42] Um, [32:43] Or just the way it uses tools is obviously very analogous to the way that a model might write unit tests and execute them and iterate over and over again until it passes them. [32:52] So I thought that was actually a really, really interesting find. [32:55] Really interesting. Did you see Taki? [32:58] No, it's the language model that's trained only on text from before 1930. Oh, yeah. What do you make of that? Because I thought it was so interesting that you can get it to you can get it to program if you if you if you if you shot prompt that you can get it to program like basic things. What do you make of that? And what does that tell you about the value of data? [33:18] So I personally didn't dig into it that much, but I thought the concept was fascinating. Like basically this idea, and I think a lot of people have this idea, it's like,
[33:27] If you gave, if somehow were able to create a data set, I think contamination issues are very, very difficult to avoid. So the question is how you would do this. It's like if you gave the model, you know, data only up until... [33:38] you know, pre-Mutant, would it be able to discover Neukonian mathematics? [33:44] Would it be able to discover quantum physics and so on and so on? [33:49] So yeah, I think it's a really, really interesting question in terms of what types of [33:54] inherent reasoning, [33:56] the model [33:58] We'll be able to learn it. [33:59] and then extrapolate from that. And it's almost like if you can discover all those things, then okay. [34:06] then given the state of science today, does that mean that the model is going to be able to discover science that centers out? [34:12] Having played with it a lot, [34:15] My sense is... [34:16] The answer is no. [34:18] but a qualified no. [34:20] and you can kind of feel it, [34:24] You can feel it bumping up against the limits of its world when you start talking to it about more modern things. [34:30] There's this foster science, Thomas Kuhn, he talks about incommensurability. And it feels like my world and its world are sort of incommensurable. But then you can also get it to... [34:42] program. But the way you do that is you get it to combine its circuits in a way that's not... It wouldn't be natural for it, but you can prompt it in a way to do that, in a way that ends up being programming. So... [34:54] I sort of both think it can't do it. And also, if you prompt it cleverly enough, it can. But you have to supply the answer first.
[35:03] Does that make sense? [35:05] Yeah. Interesting. Okay. What is the value of my data? [35:10] So one of the things that I'm just so interested in... [35:15] Obviously. [35:16] You're in a data company, like you're, you're, you're getting expert data from like real PhDs and, and selling it to the model companies and like providing all of the, all the like smarts and taste to, to the models that we use every day. [35:32] For someone like me, [35:34] uh, [35:35] We're just getting to a point where it's actually pretty easy for me to gather a data set. You know, like, for example, I do all of my email in codecs. [35:45] And [35:46] I have a history for every email of, was this useful? Did I dismiss it? Did I reply to it? If I replied, like, what did I say? Um, yeah. [35:55] What is the value of that? If I wanted to sell that to you, how much would you pay for it? [36:00] So the value to me as someone who would use that data to train an AI model? [36:08] Let me think. So... [36:11] I think the value would be... [36:15] teaching models, very, very deep personalization. [36:19] I think right now the models are actually not very good at personalizing things. It's kind of funny. Whenever I use AI models, I actually turn off the features where they personalize to me or where they can search across all of my conversation histories and, [36:32] Because I find that they just...
[36:35] over-index on things that I said once, but actually aren't all that important to me. [36:41] So I actually have it completely turned off. [36:43] Unless I'm testing something. So I think the value of it would be like, okay, yeah, you did report all of these emails as spam. So the next time this email comes in, you should automatically know that it's spam. [36:55] Or it's your learning dad. This is your writing style. Like one of the reasons I think people don't use AI for better or worse for writing more is because it sounds obviously I generated and it's not matching their, their voice or their cadence. Yeah. [37:09] Or it's that, okay, these are the things that you yourself care about. [37:12] I think one of the biggest... [37:14] One of the reasons AI is maybe not as useful as people would have expected sometimes is because it lacks all of your context. Like it doesn't know that these are the articles that you read. It doesn't know that these are the decisions about, you know, the company that you're making. These are the goals that you have. [37:32] And once... [37:35] All of that is insane. [37:37] in the model's history and it knows that it can incorporate these things and these are the kind of optimal decisions that you made. [37:45] It's very, very valuable in teaching it, okay, this is actually how I use all this data to make certain kinds of decisions. [37:51] So, yeah, I think that deep personalization is what is most unique about that. [37:57] That's interesting. And as an individual person, I mean, I guess I could turn it into a synthetic data set. [38:04] But as an individual person, is that worth a lot? Like, should I be thinking about selling it?
[38:09] I imagine we could make you an offer. I had to learn a little bit more about how big this data set size is. I mean, I can make it as big as you want. I've got Fable. [38:23] You convinced me. [38:24] One of the things we actually do is, I mean, we teach models in these very, very deep, personalized ways. [38:32] So something similar to what you described is a fairly big thing. [38:35] Tell me, tell me more. So like, I mean, I've got email, like what else, what else am I doing that you're like, oh, that's actually really valuable and important in ways that people probably wouldn't know. [38:46] So honestly, even things like the way you interact with your browser, [38:52] is interesting. Models still aren't all that good at it. [38:56] Or... [38:58] even the types of conversations that you're having with AI. [39:02] That is just inherently interesting in of itself. [39:05] Models themselves are not very good at generating synthetic conversations to try to mimic you. [39:13] And so even just knowing what types of conversations you're having is helpful. [39:17] Or it's like the combination... [39:19] It's like the combination of all these things, like knowing that these are your photos, these are your texts, these are your slacks. It's like this interconnected web. [39:27] and maybe certain things and one aspect of that web influence others. So just seeing the thing as a whole is very helpful as well. [39:35] Why are models bad at writing and how does that relate to the personalization challenge?
[39:40] So, I think some of the models are pretty good at writing, but... [39:45] Some of them are actually kind of shockingly terrible. So I'll give an example. So we created a benchmark called Hemingway Bench a couple months ago. [39:55] And it was designed to test models creative writing abilities. [39:59] And one of the things that we saw was that some of the models [40:04] They were literally outputting metaphors in every single sentence. [40:10] And I think the reason that was happening is because I've talked a little bit about this phenomenon of reward hacking. [40:17] It's almost like... [40:18] there was a metric somewhere or like a score that these models were getting. Like, okay, every time you are, [40:25] literary... [40:26] Every time you're using... [40:27] Complex imagery, it would get a point. [40:30] And it learned to reward this by... [40:34] outputting a metaphor in every single sentence. [40:38] And I mean, what's kind of funny is that a couple of weeks ago, there was this kind of like semi-prestigious literary prize, I think the Commonwealth Prize. [40:47] And there was a controversy because a clearly AI generated story won the prize. [40:53] And if you actually looked at that story, it's funny. Like it literally had a metaphor in every single sentence. And so this kind of phenomenon that we described a couple of months ago, yeah, it was still happening. [41:03] And so... [41:05] Yeah, I mean, I think it boils down to a couple reasons, but one is people are kind of sort of,
[41:11] measuring the wrong thing. Like instead of measuring actual taste and actually good prose, [41:20] they either had these flawed metrics like [41:24] What is the complexity of the prose I'm writing? How many metaphors do I have? [41:31] Or there are these AI leaderboards, again, like Ella Marina, where you have people who are essentially high schoolers who are reading responses for two seconds and, [41:40] And what they are captivated by is a flashy metaphor. [41:45] And they are not captivated by kind of like the understated pros. [41:49] And so I think it kind of boils down to a mismatch in measurement and a mismatch in the optimization and objectives that the models are trained towards. [41:59] Fascinating. Okay. Last question. What is your current AGI timeline? [42:04] So, I certainly believe that AI will happen more than most people expect. [42:11] Like every few months and even faster now, I think, [42:15] What AI is doing continues to surprise us. [42:18] So I think it depends a little bit, obviously, on your definition of AGI. [42:24] But if my metric or something like [42:27] Being able to... [42:31] automate the work of the average engineer [42:35] or being able to publish more and more novel scientific research, [42:40] that...
[42:41] gets published in these journals, or even the ability to win a Fields Medal or a Nobel Prize. [42:47] I could see it happening within the next five years. [42:52] All right. [42:53] Edwin. [42:54] Thanks so much for joining. [42:56] Thanks for having me. [43:05] Oh my gosh, folks. You absolutely, positively have to smash that like button and subscribe to AI&I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard. But instead of gold, it's filled with pure, unadulterated knowledge bombs about chat GPT. [43:27] on the edge of your seat. [43:29] craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor. Hit like, smash subscribe, and strap in for the ride of your life. [43:42] And now, without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you.
Want to learn more?
Ask about this episode