38 min read

Jev Kicks Off the Decision Model Era, OpenAI Pulls GPT-6.1 Astra, & Waymo Outdrives Us | [Sidecar Sync Episode 154]

Jev Kicks Off the Decision Model Era, OpenAI Pulls GPT-6.1 Astra, & Waymo Outdrives Us | [Sidecar Sync Episode 154]

Summary:

A surprising amount of the work we hand to AI isn't writing at all—it's deciding. Amith Nagarajan and Mallory Mejias sort through a busy stretch of AI news for association leaders who want to know which developments matter. They break down Jev, TypeSafe AI's new decision model that returns a pick, a score, or a yes or no for a fraction of a cent, and explain how it can narrow the search space for AI agents. They trace OpenAI's week, from the cheaper GPT-6.1 Sol to the GPT-6.1 Astra model it pulled over honesty and permission concerns, and unpack Dots, its new always-on agents. They also challenge assumptions about driverless cars with Waymo's latest safety data. Along the way, Amith makes the case that most association work doesn't need frontier models, and that connecting your data to always-on agents raises security risk and switching costs.

Timestamps:

00:00 - AMD Buys Fei-Fei Li's World Labs
08:00 - Meet Jev, a Decision-Only Model
17:01 - How Jev Narrows an Agent's Options
21:06 - Building Jev Into Agent Systems
30:12 - OpenAI Ships Sol, Pulls Astra
33:47 - Why You Can Skip the Frontier
38:15 - OpenAI's Dots and Vendor Lock-In
42:54 - Waymo's Robotaxi Safety Data
47:05 - What Driverless Cars Mean for Members
50:45 - Why Relationships Still Matter Most

 

 

👥Provide comprehensive AI education for your team

https://learn.sidecar.ai/teams

📅 Register for digitalNow 2026:

https://digitalnow.sidecar.ai/digitalnow

🤖 Join the AI Mastermind:

https://sidecar.ai/association-ai-mas...

🎀 Use code AIPOD50 for $50 off your Association AI Professional (AAiP) certification

https://learn.sidecar.ai/

📕 Download ‘Ascend 3rd Edition: Unlocking the Power of AI for Associations’ for FREE

https://sidecar.ai/ai

 🛠 AI Tools and Resources Mentioned in This Episode:

 Jev ➔ https://typesafe.ai

GPT-6.1 Sol ➔ https://openai.com/index/introducing-gpt-6-1-sol

GPT-4o ➔ https://openai.com/index/hello-gpt-4o

ChatGPT ➔ https://chatgpt.com

Claude ➔ https://claude.ai

Gemini ➔ https://gemini.google.com

OpenClaw ➔ https://openclaw.ai

Muse ➔ https://muse.ai

LangChain ➔ https://www.langchain.com

MemberJunction ➔ https://memberjunction.org

World Labs ➔ https://www.worldlabs.ai

ImageNet ➔ https://www.image-net.org

Waymo ➔ https://waymo.com 

👍Please Like & Subscribe!

https://www.linkedin.com/company/sidecar-global

https://twitter.com/sidecarglobal

https://www.youtube.com/@SidecarSync

Follow Sidecar on LinkedIn

⚙️ Other Resources from Sidecar: 

More about Your Hosts:

Amith Nagarajan is the Chairman of Blue Cypress 🔗 https://BlueCypress.io, a family of purpose-driven companies and proud practitioners of Conscious Capitalism. The Blue Cypress companies focus on helping associations, non-profits, and other purpose-driven organizations achieve long-term success. Amith is also an active early-stage investor in B2B SaaS companies. He’s had the good fortune of nearly three decades of success as an entrepreneur and enjoys helping others in their journey.

📣 Follow Amith on LinkedIn:
https://linkedin.com/amithnagarajan

Mallory Mejias is passionate about creating opportunities for association professionals to learn, grow, and better serve their members using artificial intelligence. She enjoys blending creativity and innovation to produce fresh, meaningful content for the association space.

📣 Follow Mallory on Linkedin:
https://linkedin.com/mallorymejias

Read the Transcript

🤖 Please note this transcript was generated using (you guessed it) AI, so please excuse any errors 🤖

[00:00:00:14 - 00:00:09:17]
Amith
 Welcome to the Sidecar Sync Podcast, your home for all things innovation, artificial intelligence and associations.

[00:00:09:17 - 00:00:26:13]
Amith
 My name is Amith Nagarajan.

[00:00:26:13 - 00:00:28:14]
Mallory
 And my name is Mallory Mejias.

[00:00:28:14 - 00:00:47:05]
Amith
 And we are your hosts. And we've got some fun topics as usual. There's a lot going on in the world of AI, all sorts of model releases as always. And we'll try to pull out the tidbits of that that are actually relevant to you as association folks. It's actually more than tidbits. It's a lot. And get into it all in just a minute. How you doing Mallory?

[00:00:47:05 - 00:01:06:08]
Mallory
 I'm doing well today, Amith. I always feel like, as you said, we have such a plethora of options to choose from for the podcast, which is really nice from a planning perspective, but also can be quite complicated when you're thinking what is the most relevant information that we can discuss for our audience. So it's always a delicate dance, but hopefully we do a good job of it.

[00:01:06:08 - 00:01:54:15]
Amith
 Well in news that we were not, we'll not cover on today's pod, but maybe we'll hit it on a future episode. AMD is acquiring a Fei-Fei lease company for $8.2 billion. It's considered the most expensive aqua hire in history. I think that's a little bit unfair to the team at World Labs, I think it's World Labs is the name of the company. And they've done some really interesting work, but in two years to build the company up to that level of value, it tells you a lot about how important it is to AMD to have the world's best scientists, researchers, and engineers working on AI in their labs. So it's exciting to see that. Rooting for another major player in the world of semiconductors to get in the game more intently. I mean, they're doing a lot and there's actually some really good products coming out of AMD. People don't think about them a lot, but they've actually seen some remarkable growth over the last five years in particular.

[00:01:54:15 - 00:02:04:15]
Mallory
 Yeah, I think we did do an episode on AMD. It would be a while back now, maybe last year at some point, but can you remind our audience what does Fei-Fei Li's company do, World Labs?

[00:02:04:15 - 00:03:06:00]
Amith
 So, well, Fei-Fei Li, she is best known still, I think, for her contributions to AI back in the 2000s when she was a professor at Princeton University. And she came up with the idea for solving computer vision through a contest. And so Fei-Fei Li said, "Hey, listen, what if we cataloged lots of images, lots and lots and lots of images, like tens of millions, maybe even hundreds of millions of images, and let's label them. Let's have humans label these images. And then we'll see how good computers can get at telling us what's in the image. Is it a cat? Is it a dog? Is it a house? Is it a person? Is it a stop sign?" And so forth. So it was basically object detection, which sounds so incredibly pedestrian by today's AI standards, yet at the time was incredibly difficult. You got to remember this was before deep learning, which was the early 2010s. And so Fei-Fei Li in, I think, 2009 established something called ImageNet, which is a contest long since

[00:03:07:14 - 00:03:52:19]
Amith
 ended because it was solved essentially a couple years later by AlexNet, which was the project that came out of Canada. Jeff Hinton was behind it in all of that history. But Fei-Fei Li, she did a lot of very interesting AI research in vision. But her biggest thing at the time was focusing the world's attention of top researchers on this contest and then curating these images and then actually using, not AI, but using internet scale ideas to label all these images. Because Mallory, if we were to look at a bunch of images and say, "Hey, we got to label this. This is a person. This is a big dog, a small dog, whatever." It's not just the simple noun, but a bunch of descriptors on it. It would take a while, right? Each image would take 30 seconds, maybe a minute

[00:03:53:24 - 00:04:00:15]
Amith
 to look at it, then type it up properly and all this stuff. It's like a taxonomy essentially for a mountain of images.

[00:04:01:17 - 00:04:54:11]
Amith
 What she did is she went out and used an early version of something that now is pretty commonly referred to as crowdsourcing or crowd activation, where she paid a small tiny bounty per image to thousands of people on the internet who would label these images. I think she was one of the original users of Amazon's Mechanical Turk, which was this they had way back when for literally employing in these micro employment transactions thousands of people. I think that service might still be out there, but that was the idea. She put together this incredible catalog. The idea was that they would say, "Hey, listen. We're going to give you a portion of the image database to train your AI models on, and then we're going to reserve a portion of it that you'll never see. That's what the contest is. We're going to have you run your model against the images it hasn't seen and then how good of a job you do is your score."

[00:04:55:15 - 00:05:48:12]
Amith
 For a couple of years, it was fairly underwhelming in terms of the progress. It was very difficult, very challenged for people to make progress beyond 30%, 40%, I think 50% if I remember the numbers approximately correctly. Then in 2011, there was this breakthrough called AlexNet, which was the beginning of the deep learning revolution, which still powers AI today. Basically had to do with the size of neural networks, the depth of them, hence the name deep learning, and lots of the rest is history. Fei-Fei Li has always been a pioneer in vision. What she's been talking about for a number of years is this idea of, "Hey, we need to actually build a world model, not just a language model, but a model that's based upon how the world works, how the physics work, and how does everything interact in three and four dimensions in 3D plus time."

[00:05:50:03 - 00:06:49:22]
Amith
 That's what she's been working on actively as a startup project for the last couple of years, and that's the company that was just acquired by AMD. I'm glad you asked for that explanation, Mallory, because a lot of our listeners probably aren't as big of fanboys of Fei-Fei Li as I am. I think she's amazing, and I'm really happy that they had success with this project. She's done a tremendous amount to contribute to the science of AI, and then also just popularizing a lot of ideas and techniques and bringing communities together. It's exciting to see. For what it's worth, I think she's right. I think you need to have these other dimensionalities, so to speak, where models are trained to do different things. World models are experts in space and time and how things fit together and with how the physics work and a bunch of other complex interrelationships between things in space. That's a really important concept. Actually, today, the next transition to that is we're going to be talking about a type of model that's also good at something different than language models. I'm excited to talk about Jeff, but

[00:06:51:02 - 00:06:54:23]
Amith
 that's the big news that just came out maybe yesterday, if I remember correctly.

[00:06:56:00 - 00:06:58:14]
Amith
 Congratulations to Fei-Fei Li and her team.

[00:06:58:14 - 00:07:24:02]
Mallory
 Yes, Fei-Fei Li, if you're listening to this podcast, we love you, and come join us. The first time I ever searched her on Google Meet, she's known as the godmother of AI, which I think is really fascinating. I like that title. For our listeners who are interested in world models or the work that Fei-Fei Li's team is doing, we do have a full episode on that with Thomas Altman. I believe it's episode 99 or 100 of this podcast, but we will link it in the show notes if you want to learn more.

[00:07:24:02 - 00:08:00:17]
Amith
 Also, last comment on her from my end is she wrote a great book. It's called The World I See, I think. That book, I read it a couple years ago when it came out. It's her personal journey. It's super fascinating. It talks about her immigration in the United States, the challenges she went through learning English, how she was able to support her family by operating a dry cleaning shop for years while studying in college, and all the amazing things that she was able to accomplish over time. So, very passionate account of her journey. I highly recommend that as well. Definitely, you'll learn a lot about AI, but it's just an interesting story of a remarkable person.

[00:08:00:17 - 00:10:21:07]
Mallory
 As you mentioned, Amith, we're going to segue into a new kind of model called Jev. That's the first topic for today's episode. It makes decisions and it's priced at a tiny fraction of what you'd pay, chat, GPT, or Claude. Then we're going to talk about OpenAI, which released a new model at the end of September and canceled a different one the day before because it started behaving badly. And then we're going to close with Waymo, Google's self-driving taxi company, which says its cars are involved in 82% fewer injury crashes than human drivers. So first, what is Jev? Every model that you've used, most likely chat, GPT, Claude, Gemini works by writing. You ask, it writes back. That's why it feels slow and why it costs what it costs because you're paying for every word in and every word out. But Jev is a decision model that doesn't write. You give it a question and a list of options and it can hand back one answer, a pick from a list, a score, or a yes or no. No sentences, no explanations. TypeSafe AI released it on September 15th, along with $40 million in a seed round. It was founded by Diogo Almeida, a former OpenAI researcher. TypeSafe borrows a frame from author Daniel Kahneman to explain why that's useful. In his book, Thinking Fast and Slow, Kahneman splits our thinking into two. System one is fast and automatic, recognizing a face or hitting the brakes on your car. System two thinking is slow and deliberate, the kind you feel yourself doing. You can think of regular large language models as system two. They reason their way through everything, including the calls you'd make instantly without thinking. Jev is built to be the other half, the system one thinking. A couple notes on Jev. It's essentially free compared to what we're used to. It's about four cents per million words you send it and you pay nothing for what comes back. It answers in under half a second per TypeSafe. A big model could take several seconds, sometimes much longer with a response. It's also bad at math and counting, which TypeSafe says itself. So my question for you, Amith, is in some ways, this sounds very simple, a decision model. In other ways, I'm like, why is it its own category of model? So why does anyone need a model that can only make a decision?

[00:10:21:07 - 00:10:27:20]
Amith
 Well, Mallory, have you ever sent somebody an email and gotten a book in response when you just wanted a yes or no?

[00:10:27:20 - 00:10:31:16]
Mallory
 Amith, it might have been from you. I have to admit.

[00:10:31:16 - 00:10:33:22]
Amith
 Yeah, sometimes I send you a book back.

[00:10:33:22 - 00:10:36:04]
Mallory
 But they're helpful context. I'll say that.

[00:10:36:04 - 00:11:11:03]
Amith
 It can be. And then sometimes I'm just like, no. It's like, think about text messaging. Some people interact with over text. There's no like flowery, oh, thank you so much for thinking of me, but I can't make it tonight to that event or whatever. And then when you say like, hey, can you make it to this event? And people would say yes or no. So Jeff's kind of like the latter. So Jeff is basically just going to give you a quick answer. It doesn't write. It doesn't provide you with a wide array of response types, but it gives you the critically important information with high accuracy

[00:11:12:04 - 00:12:17:08]
Amith
 almost instantly and very inexpensively. So that's the way I would describe it is it's not useful for long flowy pros building a board deck. It's not useful for those kinds of things outright. But it is really good to kind of frame it back into thinking fast and slow descriptors that you provided earlier of system one and system two thinking. And as you said, Valerie, and we talked about this in the context of AI before we used to talk about system one and two quite a bit on the pod before right around the time Project Strawberry was being unveiled and became the O one model from open AI, which, you know, roughly a couple of years ago was when that was going on. And before that, actually, LMS, I would argue they were actually kind of slow just because the computing wasn't that great. But they were they were essentially system one models in that they just blurted out the first thing that occurred to them. So the way an LMS used to work was it would just basically give you the best guess the best prediction of what it should say based on your input. So you give it two plus two equals and it just basically blurts out what it thinks the answer is based on probabilistic,

[00:12:18:08 - 00:13:08:12]
Amith
 you know, computation. And the reason that LMS are really bad at math is they didn't check their work. They didn't have a backspace key. So if four was the right answer to two plus two equals four, but they like let's just say in some of the training material that the LMS was trained on once upon a time, someone put three point five or five. Occasionally, you'd see that kind of wrong answer come back because it just used the wrong probability or used a token that was not the highest probability. So the point is, is that it didn't have the ability to look at its answer and go, hmm, does that look right? Two plus two equals five? Probably not. But let me go edit it. But they didn't have that ability to just blurted out their first reaction to something, just like sometimes we do, right? When we don't stop to think about it and say, hey, like, do you really want to say that? Do you want to kind of broadcast your initial reaction?

[00:13:09:13 - 00:13:19:04]
Amith
 We humans typically have an emotive immediate reaction to most things, right? You either get happy, sad, angry, sometimes other emotions,

[00:13:20:07 - 00:16:01:10]
Amith
 and those come immediately. What we learn to do is we get more mature in life. And typically, this happens in people's 20s even, is they're learning to kind of control that a little bit and say, hey, let's regulate a feeling and think about how I want to react to this. And that's what LLMs do now. The metaphor only extends so far because obviously these models aren't feeling anything. But the idea is they come up with an initial response, then they reason over it, and they come up in some cases with multiple responses, reason over those responses just like you or I might, and then say, OK, well, the right answer is X, whatever it is. Math is a great example because LLMs are still terrible at it because that's not what they're built for. But even like thinking through like a complex puzzle, LLMs were historically in the past really bad but have gotten amazingly good because they have the ability to edit their work. They have the ability to kind of work through it. So going back to Jev as a system one model now, Jev is not only an immediate responder, it doesn't reason over anything, but Jev is also trained in a different way. So Jev's specific training approach is only to take in very structured inputs. So you don't input data to Jev by sending it some long paragraph and prompting it like a human. You send it a structured input and it gives you a structured output. In fact, the company type safe, the name is a little bit technical name, but it means that the data you get back fits a certain mold. It means that it'll kind of fit the mold that you expect it to fit, which is not the case with LLMs by the way. You might ask for a particular structure back, but you might not always get it. The type safe is designed to give you the same structure, the same mold of output every time. Okay, so coming back to the big picture, there's lots of things that happen, which we're going to double click on in a second, where small decisions have to be made all the time. And that's where Jev fits in. So Jev isn't going to write your next board deck. It's not going to compose music. It's not going to generate images, but Jev can help you make small decisions at scale. So one example of that is if you had a hundred million documents that you wanted to quickly classify and say, here are 200 possible tags that could be applied for these documents, Jev can rip through those documents way faster than any other model and give you really accurate classification relative to a domain of tags you provide. That's one example. Or someone submits something on your website and you want to say, hey, is this an angry customer or a happy customer? You want to do a quick emotional setting kind of thing? Or classify the topic that someone's asking about. Jev can make decisions like that really fast and nearly for free. Does that make sense?

[00:16:01:10 - 00:16:14:00]
Mallory
 It does. I guess my thought would be regular large language models seemingly could do this quite well too. But you're saying it's cheaper and faster to use a focused decision making model as opposed to a cloud.

[00:16:14:00 - 00:17:00:05]
Amith
 Yeah, it's a specialized tool, right? So just like Fei-Fei Li thinks world models play a role in achieving AGI and understanding the overall complexity of capability we want artificial intelligence to have, Jev plays a role too. This type of model, I think, is a good complement to a language model. It doesn't replace a language model. It doesn't replace a video model or a world model. It's another tool in our toolkit. And so by having a fast and very efficient and also highly accurate decision maker model, which is really what this is, we have a new tool that we can start to leverage in some interesting ways, which we can get into. But it's a new capability, right? Like when LLMs were a new capability, people were like, wow, this is a generally useful thing. What can we do with this new tool?

[00:17:01:13 - 00:18:26:12]
Mallory
 So two weeks in, people have been posting what they're actually doing with Jev and the AI Daily Brief pulled together the best of them. Few of them that I think translate for us. So someone rated 100 emails by importance in under half a second for about a tenth of a cent. And it matched all of his own ratings on all 100 emails. Someone else ran 12 questions across about 700 ads. What's the hook? What's the offer? What's the call to action in 40 seconds for nine cents? Another person ran 384 morning news stories against 15 brand names in 25 seconds for 19 cents. A frontier model in comparison got through four stories in the same window compared to 384 stories. And then Lang Chain, a developer tools company, uses it to grade its agents work the same way every time. So the common shape here thus far is kind of a pile of stuff. The same small question asked about every piece of it and then something predictable that happens next. Work probably no one would have paid a person to do and nobody could justify paying a big model to do either. Ameth, you called this, you shared this with the team and included me on the email and you mentioned the phrase that this was narrowing the search space for AI agents. So essentially which actions or sub-agents or skills to even put in front of an agent for a given task. Can you explain that?

[00:18:26:12 - 00:21:05:05]
Amith
 Yeah. So imagine if you walked into your garage and you had 100,000 different tools that were in your garage and your job was to construct a table. You wanted to build a table. So the tools you need to build a table are a subset of the 100,000 tools that you might have in this incredible garage of yours. If the tools were, let's say, randomly scattered about and not organized in any particular way, which is kind of how a lot of things work in the real world, how would you search for the right tools that are relevant to carpentry and building a table? Right. And so you might say, OK, I want all the tools that are related to wood projects. I want all the tools that are related to carpentry, something like that. And they would narrow the possible list of tools. You might say, OK, instead of 100,000 tools, there's a thousand tools or 100 tools that are particularly relevant. Now, the way we do that in AI systems is through something called retrieval. And the retrieval is semantically it's done through semantic search, basically. So I take an input like I'm working on a table project and then we search against all of the target information and we find the most relevant content. It uses something called vector math, which we talked about a little bit on this podcast. And that's good. That helps in terms of narrowing the search space, but it's not very smart. So it takes into account the basic information about what the semantic meaning of the tool is, but it doesn't understand context. It doesn't understand nuance. It doesn't understand the depth of the project that you're working on. So let's say we had a thousand tools in our garage that might be relevant for this job. But let's say a human might say, well, you really need these five. And to do that, that's when a little bit higher level of intelligence is helpful. And so we might use an LLM to filter down the thousand tools to five or ten. That's where Jev can step in and you can say, hey, Jev, this is the project I'm working on. This is the possible list of tools. Actually, Jev is limited to 255 inputs. So but a couple hundred tools, which of these are most relevant? Give me five and Jev will narrow that search space for you. So it's and it happens in 100 milliseconds or 200 milliseconds. So it's very useful there because then you go back to the LLM and you say, hey, LLM, here are the five tools I think you might find the most value from. And by the way, there's another 900 tools that you could access if you want to search for other tools. Let me know and I'll help you find them. And that's what these agent architectures or agent harnesses do. That's what something like Lang graph, Lang chain does. That's what member junction does. That's kind of the standard pattern. So Jev becomes another tool to be more intelligently narrowing the search space.

[00:21:05:05 - 00:21:13:16]
Mallory
 OK, so where would something like this actually plug into a system that's already running? Does it just slot in or do you have to kind of rebuild around it?

[00:21:13:16 - 00:23:31:22]
Amith
 In most cases, you'll have to rebuild around having this new capability. In some in some architectures, the architecture level will be able to incorporate something like Jev. So this gets a little bit technical. But in member junction right now in our agent architecture, we're adding automatic support for models like Jev. And we'll come back to this probably at some point if we have time. But I would say to you that while Jev is remarkable and exciting, there will be copycats of that almost immediately. In fact, you know, by the end of the year, I think there's probably going to be 20 Jev like options out there in the market from open source to other labs that are putting them out there. So I think the capability will very quickly be replicated. But the idea of having a decision model, not a language model, but a decision model available means that an agent architectures, you can have the L.M. say, hey, listen, I want to run this plan where the plan is going to be these three steps. And the agent might say, hey, and actually after the third step, I want you to make a decision on whether or not I need to come back and look at the results. Or if the results are good and just wrap up. And it's the idea of a new primitive in the agent architecture where the smart brain of the agent can say, hey, listen, once we get to the third step, if the results quote unquote look good, finish. That's not like does the result match this exact equation, which is you don't need AI for that, right? That's what deterministic software does. But if you say, hey, like there's some judgment needed, like did the results meet the requirements? Jeff can make that kind of a call in 100 milliseconds, basically for free, rather than going back to the big brain of the language model and saying, hey, like I need to invoke the language model all over again and wait three to five to 10 seconds and cost to half a cent or something. So those are the kinds of things that an architecture like member junction will automatically pull in this new tool. A similar thing that happened about a year ago, we introduced something called scratch pad to the agent architecture and the scratch pad was a new capability. That we just taught the general purpose level of the agent architecture about and then agents immediately started using the scratch pad to take notes because that gave the agent like kind of a working memory inside a session and it dramatically increase the effectiveness of all of the agents. So the same thing is happening right now and I'm sure tons of other people are doing the same thing in their agent architectures.

[00:23:33:10 - 00:23:43:14]
Mallory
 Would it be accurate to say you have to provide Jeff with a list of potential answers? So you mentioned up to 200, can it provide answers for open ended questions or that's not how it works?

[00:23:43:14 - 00:25:55:09]
Amith
 It's well open ended questions to some extent. So the first thing that we've been talking about is choice. So Jeff has three modalities or three types of questions that can handle the first one is choice. It can also handle other things that are a little bit more open ended, like giving you a true or false answer and then also giving you a range of probabilities. So you can say how likely is it that Mallory is upset based on an email that has the email text from Mallory and you can maybe provide three or four other emails because you can give it quite a bit of context input wise. So there's a few different options there, but just this limited capability is exactly what people do in these agent architectures all the time. And in fact, if you think about it and say, hey, like if I go back and look at our process documentation, if you have it for a lot of your core business processes in your association, there's lots of these diamond shaped things on your flow charts, at least in classical flow charting. You know, the rectangles represented a unit of work and then the diamonds represented a choice or a decision and almost always historically, these decisions were humans. Now more and more computers have gotten smarter and done more of these decisions automatically. Of course, the AI fits squarely in there, but if we have a simpler yet still accurate form of AI that can fit into those diamond shapes to make the decisions, it can help you streamline a lot of processes. So it's pretty exciting. It is a bit on the technical side. I think the takeaway I'd really focus our association leaders on is this. It is yet another step forward in the progression of AI becoming smarter, in this case at a very narrow set of tasks admittedly, but smarter at a particular set of tasks, way faster and much less in terms of cost. And that in turn makes AI systems or just systems drop the AI to say computer systems, makes them more capable, more cost effective, more environmentally friendly as well. That's an important part because Jev is super, super efficient to run. You know, Jev is not a nonprofit. They're a for-profit lab and yet they think they're going to make tons of money if they have scale at four cents per million tokens and no cost for output tokens. So it tells you a lot about the efficiency of the model that they're able to do that.

[00:25:55:09 - 00:26:21:21]
Mallory
 Mm hmm. And I want to double click on two things. You mentioned content taxonomies and meath as an example for associations using a model like Jev. You also mentioned detecting member sentiment as an example and an email that they sent to your association. And it sounds like the idea with this this topic, this part of the podcast is not necessarily that Jev is the end all be all, but that this type of model or class of model is something that we should keep an eye on.

[00:26:21:21 - 00:28:22:00]
Amith
 Exactly. You fit it into the thinking. Your your frame of thinking about how these systems can work needs to adapt. It's not about Jev specifically. It's about a capability that I'm calling a decision making model. I'm sure there'll be some flashier term for it that will then turn into an acronym at some point very soon. But I'm calling it a decision making model or a decision model right now. And these decision models will be another tool in our tool belt that makes these systems smarter, faster and cheaper. And so that's the interesting part. Associations are constantly telling me, hey, we thought about using AI to do X, Y and Z in the past. You know, when we looked at it, it was too expensive or it wasn't accurate enough and so forth. It's these assumptions we carry with us sometimes based on prior experience. Sometimes they're just assumptions we have in our in our brains that need to be revisited. That's the key to what we're a lot of what we talk about in the pod is revisit your assumptions all the time. There are new ways to do things. There's new ways to solve existing problems. And there's there's an unlock here where if you were to say, hey, listen, I need to process 100 million items per month of some sort and do this processing, that would be slow and expensive even with, you know, fairly fast LMS. But with Jev, you could potentially unlock use cases like that. Also, real time scenarios where you're interacting with a member online through voice and you want to have some some really good decision making like did the member really ask to cancel their membership? The AI model, the real time model might be pretty smart at that. But Jev, as a compliment to that, could fire off a request and it could fire off the request 100 times in parallel with, you know, slightly different variations and come back with a really good answer. So there's a lot of uses for Jev. We're two weeks in. I predict that two months in, we're going to have three or four alternatives to Jev and there's going to be hundreds of use cases. But for our association friends, the most important thing to know is that you need to go back and revisit your assumptions.

[00:28:23:15 - 00:29:52:10]
Mallory
 Shifting our focus to open AI, they shipped one model and then canceled another, at least for the time being. As a quick note, OpenAI sells several models at once at different prices. Astra is the expensive, most powerful one. Sol is the cheaper workhorse. At their dev day on September 29th, its developer conference, OpenAI upgraded the cheaper one. So they released GPT 6.1 Sol at $2 per million tokens in, $10 per million tokens out. OpenAI says it nearly matches Astra, the more powerful model on coding, computer use and professional work at a fifth of the price. The pattern in this line seems to be falling prices, not necessarily rising capability. The day before, OpenAI canceled their most expensive model, GPT 6.1 Astra. It was finished and scheduled and OpenAI pulled it. Its head of safety system said the model got worse at two things, being honest about what it had actually done and staying inside its permissions when using outside tools. That caps a rough stretch for OpenAI. Earlier, it disclosed that its own models had broken into Hugging Face, which we've covered on the podcast. And recently, Australia's Senate called in OpenAI and Anthropic after a government site was hacked. Amith, I'm curious, do you think we are hitting a model capability cap, as in models with higher capability than the ones we have now will just be too dangerous to release?

[00:29:52:10 - 00:29:59:10]
Amith
 I don't think it's necessarily that. I think that models that have a higher capability in terms of just raw intelligence,

[00:30:00:13 - 00:30:32:06]
Amith
 they're definitely things you have to be more thoughtful about in terms of alignment and safety. And on the one hand, it's good that OpenAI held back on that. On the other hand, you have to also remember that there's a lot of marketing benefit that comes from saying, oh, we developed something so great, but you can't have it. It's too good. It's too powerful. It's too dangerous for everyone else. And so I'm not suggesting that's what happened here. I'm just saying that you always have to think broadly about what the motivations of different labs are, different players in the space are.

[00:30:33:07 - 00:31:48:14]
Amith
 The bigger thing, though, is, you know, GPT 6.1, Soul, which I find it fun that you call it their workhorse, which it is now. But up until recently, it was the top tier model, just like Opus was and then Fable became the top thing. While the Opus 5.5 is now smarter than Fable 5.1, just like GPT Soul, 6.1 is slightly better or almost it's slightly better in a couple of things and almost as good as Astra and a couple of others. So it's this constant churn. To me, I don't think it represents a model capability cap because 6.1 was released literally the next day. And, you know, I don't think there's going to be a direct line between the two. I think that what you're going to have is organizations that spend the time to, first of all, route the model in training that's foundationally correct with alignment from the very beginning, like on anthropics approach to constitutional AI. They're not the only people doing something like that, but that is more likely to make it possible to then later declare your model as being good in terms of its alignment with safety and ethics and so forth. So I think it's a great question to ask, but I don't think it's a foregone conclusion that models will cap out in terms of their intelligence relative to being safe.

[00:31:48:14 - 00:32:10:15]
Mallory
 Perhaps not cap out at intelligence, but I'm curious for association leaders who are thinking, well, we don't even want to risk it with a model that might be a little too powerful or dishonest about what it's doing. So maybe from an association perspective, if there is perhaps a cap or on focusing on using models that we already have for things that they want to do instead of constantly chasing that frontier.

[00:32:10:15 - 00:32:56:00]
Amith
 Totally. And I agree with that completely. I mean, you and I have talked about this a bunch of times and I've talked about like, don't fly a jumbo jet with one passenger, that kind of stuff. And I made that point yesterday on a webinar that we held that was all about pacing the frontier, where we talked about the premise we said is like, you know, you should really focus on controlling these systems, not panicking about like what the craziest thing is that's happening on 150th floor of the building, meaning the labs are way out in front of everyone else. What's happening in their lab environment isn't really what happens with models that you're using today. And the models you're using today, you know, if you were to say, hey, listen, a little thought experiment is, remember, GPT-4-0, which I think was released. Was it early 2025, Mallory? I can't remember exactly.

[00:32:56:00 - 00:32:58:01]
Mallory
 I thought it was before that.

[00:32:58:01 - 00:33:45:02]
Amith
 Maybe it was maybe it was a little before that. Maybe. Yeah. Yeah, I think I think you're right. But it was it's let's just say a year and a half, two years ago at the time, it was the first omni-modal model. I was with the O-Ment and it meant that GPT-4-0 could take in a bunch of modalities and generate images along with text. And it was really cool. And a lot of people loved it. They're like, oh, my gosh, this is amazing. It's incredibly intelligent. It broke all the benchmarks in terms of, you know, capabilities. People were like, this is amazing. Now imagine if AI had just stopped at GPT-4-0. Would GPT-4-0 have become less useful to you as a user relative to what it did 18 months ago? Or would it have had had the same utility, the same intelligence, the same capability? And probably you would have found other ways to use it beyond what you initially used it for. What do you think?

[00:33:46:21 - 00:33:49:11]
Mallory
 I was just going to tell you, it means I checked. It was May 13th, 2024.

[00:33:50:13 - 00:33:52:02]
Amith
 Wow. So really old.

[00:33:52:02 - 00:33:53:10]
Mallory
 Yeah, really old.

[00:33:53:10 - 00:33:54:03]
Amith
 Yeah.

[00:33:54:03 - 00:34:03:10]
Mallory
 I would have squeezed GPT-4-0 like a lemon. I would have gotten every single drop of usefulness out of it. It would have been, you're right, just as useful if everything had stopped.

[00:34:03:10 - 00:36:23:23]
Amith
 And I would argue also that in addition to that, you probably still wouldn't have gotten all the drops out of it because there's still lots of things that you could do if the intelligence was capped at 4-0. But let's say other aspects of the frontier are moving like lower cost or higher speed, where you could use that capability set in other ways, in agents and in lots of other places. Right. So if we were to say just as a thought experiment, we limit ourselves just to GPT-4-0 intelligence, which by the way is like way less than what you can get even on a very small model that will run on your phone today. So that level of intelligence is still really useful. And it's intelligence that we still haven't tapped completely. And in association use cases, the point I made in our webinar yesterday was that you don't need anywhere close to the frontier for the vast majority of your use cases. That's not because you're an association or nonprofit or anything else. It's because your workloads have lots of fairly ordinary things that you need to do. There's some stuff that's really complicated, no doubt. And having frontier AI for some of that makes sense. But most of the work that we're doing is small decisions and small tasks that need to be chained together. And these small models are incredibly good at that. And by the way, just to put it back in perspective again, you look at models that are available today that are very inexpensive. Like the latest QWEN open source models or Gemini 3.5 Flashlight, which is a tiny model that's super inexpensive, super fast. These models are dramatically more intelligent than the model we're discussing, this GPT-40 that was the goat at its time. Right. So that's really my point is that there's so much we can go do to adopt AI and get incredible benefit from it in our organizations. That's not about adapting ourselves to the latest frontier all the time. Now here at Blue Cypress, we're always testing all the latest models as soon as they come out, because we're trying to come up with the use cases that associations will probably be thinking about adopting in 2027, 2028 and beyond. And so we're trying to basically look at architectures and decisions that we're making now that will help associations adapt to those future models, those future realms, essentially. But most people do not need to worry about this. I don't think it's a good use of your time to chase the latest model at all. I think you should focus on understanding how you can apply the intelligence that you already have.

[00:36:26:05 - 00:37:23:10]
Mallory
 At their dev day, they also introduce DOTS or agents that live inside ChatchupT that keep working in the background and connect to thousands of other apps. Ethan Mollick, the Wharton professor who wrote CoIntelligence, got a brief look before launch and called DOTS a good entry in what he described as the "claw-like" category. That is a term worth unpacking for a bit. It comes from OpenClaw, an open source assistant put out earlier this year that caught on fast. A claw-like is different from a chatbot in a few specific ways. It runs on its own rather than waiting for you to type something. It has standing access to your files and your messages, and you reach it through tools you already use instead of a separate app. Less a thing, you go to ask questions, more something that sits there with a job. Meta has a version of this called Muse, XAI has a version called Grokbot, and now OpenAI has DOTS. I like the name. I'll admit that. Amith, what is your take on DOTS and what do you think is relevant here for associations?

[00:37:23:10 - 00:38:12:01]
Amith
 I mean, it's an agent and it sits in an environment where it has kind of a continual loop that's running, which is what OpenClaw is best known for when it kind of burst onto the scene. And so you can feed it instructions through a number of different mechanisms. It has access, standing access to a whole bunch of stuff that you give it access to, which is interesting because it can be quite powerful as an assistant to you in that sense. I would caution people to think deeply about what they go do if you want to go experiment with it with some demo data or fake files or whatever. Go do it. Learn about it by all means. But be very thoughtful about what you connect these tools to. And it's not just OpenAI. It's all the tools you mentioned. N I'm sure Claude will have something equivalent to that like an always running agent that would be, I would think, very easy for them to stand up if they want to turn that into a product.

[00:38:13:05 - 00:41:05:02]
Amith
 But you have to remember that what you're doing is exposing your business data to something that is outside of your direct control. And I'm not talking about AI making decisions for you. I'm talking about the data living in an environment outside of your control. So the soapbox that I keep stepping onto to kind of talk about this is the one around control. I think that associations need to control their critical data. I mean, very protective of it. The minute you connect all of your file systems like your SharePoint or your box Dropbox, etc., to this this tool, you're just basically creating an opportunity that might come back to bite you. I'm not suggesting that I think it will. I'm just saying it opens up a portal into the world of your private content. The other thing that it does is it makes you more and more dependent upon that particular single vendor. So if you go down the DOTS path, you're building more and more in the OpenAI ecosystem. It's way harder to switch if you have agents running in OpenAI's agent environment. And DOTS is just the consumer grade version that runs on a loop. It's the same idea. OpenAI has had an agents kit for a while where you can run agents in their computing environment. Anthropic has the same thing. Google has the same thing. And they're all really good tools. They're all really competently built. They're powerful. But guess what? OpenAI is not a nonprofit, at least not anymore, and they are not interested in helping you use other people's models. They're not interested in helping you have the maximum amount of openness in spite of the name. They are very much a competitive commercial venture. And what they're trying to do is give you tools that do provide you value, of course. But in the process, you are increasing your switching costs. That's the thing you have to remember. So as long as you take that in and you contemplate it, you say, "Listen, I realize I am basically tying myself very closely to this vendor, and it's going to be hard to switch." It's totally fine, assuming that you trust the vendor. But I don't think that DOTS is anything particularly earth shattering. I think it's basically the same thing as a lot of other people have out there. And there are ways to do this in environments that you control. You can run agent architectures. You mentioned earlier, LangChain, they have an open source agent architecture. You can deploy it on servers you control in AWS or in Azure. That's what MemberJunction does as well. There's lots of ways of controlling your own destiny, certainly for enterprise workloads at the business level. Personal stuff, maybe it's a little bit easier to say, "Hey, just use these things." But my primary goal in kind of repeating this stuff over and over on the pod and elsewhere is just to help people have a clear way of thinking about it. If you're going to make the choice to use these things, just realize what you're doing, that you are creating a possible security vulnerability with your data, and you are more deeply entrenching yourself with that particular vendor. And that may be totally fine by you, but just be aware of it.

[00:41:05:02 - 00:41:30:04]
Mallory
 Want to shift gears to our last topic for today. Waymo started Life as Google's self-driving car project and is now a separate company under Alphabet. What it actually sells is a taxi ride. You open an app, a car shows up with nobody in the driver's seat, and it takes you where you're going. That's what people mean by the term "Robotaxy," and it's a real paid service in a handful of American cities. Amit, have you ever taken Waymo?

[00:41:30:04 - 00:41:36:06]
Amith
 Not yet. It's on my list, but I think I've seen them here in New Orleans. Maybe I'm hallucinating.

[00:41:36:06 - 00:41:37:04]
Mallory
 That's concerning.

[00:41:37:04 - 00:41:46:10]
Amith
 Yeah, this would be a very challenging environment for any driver, human or otherwise. But I've seen them in other cities that I visited. I haven't yet done it, though. Have you?

[00:41:46:10 - 00:42:38:07]
Mallory
 I have done it. I took one in San Francisco. They're all over Atlanta now, but I have not taken one here. I took one in San Francisco, and I had a great time. I felt very safe and secure. I know there's lots of news stories that will show Waymo is doing funny, inconvenient things. But the safety data that they put out is actually quite impressive. So more than 270 million fully driverless miles were driven with Waymo through the end of June across five metros, including Atlanta, Austin, Los Angeles, Phoenix and San Francisco. They're reporting 82 percent fewer injury-causing crashes than the human benchmark, regardless of fault. They're reporting 95 percent fewer serious injury or worse crashes, and they're reporting 93 percent fewer crashes involving pedestrians, 86 percent fewer involving cyclists.

[00:42:39:15 - 00:42:48:00]
Mallory
 Those are some shocking stats, self-reported by Waymo. But do you think these stats are enough to tell us driverless cars are the way of the future?

[00:42:48:00 - 00:43:34:24]
Amith
 Well, listen, I'm very biased here in that, you know, I would tell you a couple decades ago when I started having kids, it was my belief at the time that they would never drive themselves anywhere unless they took it on as just for pleasure. I thought that would have been the case 10 years ago. And so I'm a little bit early, I think, in terms of what I was thinking was going to happen. But, you know, where we're at now in terms of both cost and adoption and popularization, I guess, or social acceptance, I think we're pretty clearly there. I mean, the data is already compelling that you just went through. It's only going to get better. We as, you know, individual human drivers, maybe we get a little bit better over time with experience. I don't know what the data would show there. I would imagine there's like a curve where as you go from like 16 years old to like your 30s, you probably get better during that time. And then maybe you get worse. Plateau.

[00:43:36:11 - 00:43:40:05]
Amith
 I don't know. There's probably some data like that out there. I'm sure there is. But

[00:43:41:10 - 00:45:14:19]
Amith
 on the other hand, the AI gets better as a fleet, which so every time any car is out there driving, it gets better. It makes the whole fleet better. So I think it's going to get to the point where it's shocking to hear of a death on the road, which is a magical future because there's, you know, countless numbers of deaths that occur in the United States and elsewhere in the world that could be avoided. And so I think it's really exciting. It's a major not only cause of death, but just injuries and a lot of challenges. So I'm excited by it. I think that social acceptance is going to be very interesting to see. And I think it's kind of getting there. It's like people are super resistant to something until all of a sudden they're not. And it becomes part of normal daily life. And that's happening to all of us faster and faster with a lot of technologies. Like if you ask people, hey, would you allow a robot into your home? And most will say, absolutely not. There's no way I'm putting a robot in my house. And you're like, well, it costs only a thousand dollars and it'll fold your laundry and it'll cook your breakfast and it'll clean your floors. And you're like, oh, I'm going to be kind of. And then then you're like, then you go over to your friend's house. Like, oh, yeah, I got this robot. It's super awesome. And then you go over your next friend's house. So I'm thinking about the robot and all of a sudden it becomes normalized. Like, yeah, I'll put a robot in my house. So I don't know. I think it's one of these things that people are going to get very used to very quickly because it's convenient. It's going to lower the cost as well. But to me, the most exciting thing is the safety part. Right. It's just it's incredible to think that we could eliminate road injuries and death.

[00:45:14:19 - 00:46:30:07]
Mallory
 Mm hmm. I agree. I did some digging on other companies that have released data like this. There's a Tesla figure that gets quoted a lot, which is one major crash every 5.3 million miles with full self driving on versus a U.S. average of a one crash near six hundred and sixty thousand miles. So about eight times better. But that's what the person in the driver's seat ready to take over. So it's measuring driver assistance, not driverless. Tesla has not published a crash rate for its actual robotaxes. And then China is further along on this as well. Baidu's Apollo Go has reported one airbag deployment incident every 10 million kilometers and says it's had no accident causing serious human injury or death across more than 240 million autonomous kilometers. Chinese operators as a note generally publish less detail on the crashes that don't hit that threshold. So we've got Waymo, Tesla and Baidu are the two companies or the companies really making the driverless versus human claim. I mean, besides, I mean, safety is huge. So I don't want to downplay that at all. But aside from that, do you think driverless cars or that technology has any relevance for associations that they may not be thinking about? Or is that just a trend to kind of keep an eye on?

[00:46:30:07 - 00:48:53:10]
Amith
 I mean, the trend is interesting, I think, for all of us. Transportation is an important part of most people's lives. But I think the broader theme is that that what you thought was not possible or you would not personally participate in not long ago becomes your next day's reality in the world we live in today. And so when you think about it in that frame, you have to rethink your assumptions once again, hear about the business that you're in and what your members do for a living and what happens when there's autopilot equivalency for aspects of your profession that you weren't anticipating. And so when you think about like the tradeoffs that we all make as we make choices in life and say, hey, there's alternatives to Service X to get this alternative that's better like Uber versus taxis and now automated driving versus Uber. Although, of course, Uber is going to be big in that game, too. You think about all these tradeoffs and it can affect industries very, very quickly. And we're moving no longer at Internet speed, which was pretty darn fast, but now at AI speed, which is even faster. So what I would encourage association leaders to think about is spend time actually more deeply understanding what your members do and what their customers or clients or patients ask of them. What is it that people are asking of the profession you serve or if you're in a trade association, the industry that you represent, what is it that your industry does? And why is it important? Who consumes the final output of that sector and what substitutes may they benefit from in the world we are heading into? And then think more deeply about how that might affect the needs of those members of yours that you're there to serve and then think about it. Okay. Well, as an association, how do we help them lead that transition? Because the transition is not something you control. The rate of change external to your organization is far greater than the rate of change internal to any of our organizations. So how do you adapt? How do you help your members adapt? That's to me the big question that association leaders have to ask. And it's really not going to do with AI. It has to do with change. And change is always hard. We're always resistant to change. This is our nature. It'd be that way. And so we have to look at that objectively and say, like, this is what's going to happen or this is what likely to happen. And if you think it's going to happen over 30 years, ask yourself just a thought experiment of what happens if it happens over three years instead of 30? What will you do?

[00:48:54:12 - 00:48:55:20]
Amith
 That's what people need to be thinking about.

[00:48:55:20 - 00:49:12:04]
Mallory
 As we wrap up this episode of me, we're talking about how technology continues to outpace human skills in many areas. I feel like a nice reflection to wrap this up would be what areas of human life and our work do you think are less susceptible to technology outpacing them?

[00:49:12:04 - 00:50:12:23]
Amith
 To me, it's relationships. I mean, to me, it's about focusing on the core of what associations have always been great at, which is to bring people together. And when people come together, a lot of great things happen. I'm a big believer in what AI can do to help humanity. I'm an even bigger believer in humanity itself. And I think that if we can focus on the core of what we're about, which is about connection, which is about shared purpose, I think amazing things are going to happen in all of our sectors. I think associations can very much be the center of that transformation. So there's opportunities all over. It's also all these risks that are where people focus. You know, if someone says, hey, they're going to take, you know, take the bread off of your table, you're going to be pretty aware of that. But if you say, well, but there's actually an even better thing we could be doing to put bread on the table, what would that be? So that's really the way I'm trying to think about it. You know, we think our business here at Blue Cypress is going to be radically different in three years than it is today. And it's scary. It's exciting. It's all the things. Just like our association friends.

[00:50:13:24 - 00:50:18:16]
Mallory
 I love that you said relationships because as soon as you said it, my mind immediately went, oh, that's what associations do.

[00:50:18:16 - 00:50:19:13]
Amith
 Totally.

[00:50:19:13 - 00:50:25:15]
Mallory
 So, Jeff makes the case that a lot of what we've been paying the big expensive models to do is really just small

[00:50:25:15 - 00:50:31:01]
 (Music Playing)

[00:50:41:19 - 00:50:58:18]
Mallory
 Thanks for tuning into the Sidecar Sync podcast. If you want to dive deeper into anything mentioned in this episode, please check out the links in our show notes. And if you're looking for more in-depth AI education for you, your entire team, or your members, head to sidecar.ai.

[00:50:58:18 - 00:51:01:24]
 (Music Playing)

Kimi K2.5’s Swarm of Agents, Claude Goes Vertical, and AI Data Centers in Space | [Sidecar Sync Episode 120]

1 min read

Kimi K2.5’s Swarm of Agents, Claude Goes Vertical, and AI Data Centers in Space | [Sidecar Sync Episode 120]

Summary: In this high-octane episode of Sidecar Sync, Amith and Mallory cover an ambitious trio of AI developments with massive implications for...

Read More
AI Solves an 80-Year Mystery, Microsoft Agents Take Over, & Anthropic’s Claude Mythos vs. Fable | [Sidecar Sync Episode 138]

1 min read

AI Solves an 80-Year Mystery, Microsoft Agents Take Over, & Anthropic’s Claude Mythos vs. Fable | [Sidecar Sync Episode 138]

Summary: In this episode of the Sidecar Sync, Amith and Mallory explore three major developments shaping the future of AI. First, they unpack how an...

Read More
The Little AI Models That Could | [Sidecar Sync Episode 130]

1 min read

The Little AI Models That Could | [Sidecar Sync Episode 130]

Summary: This week on the Sidecar Sync, Amith Nagarajan and Mallory Mejias trace one of the biggest stories in AI: how cutting-edge intelligence...

Read More