42 min read

Rogue Agents, the Open-Weight Revolt, & Claude Gets Cheaper | [Sidecar Sync Episode 145]

Rogue Agents, the Open-Weight Revolt, & Claude Gets Cheaper | [Sidecar Sync Episode 145]

Summary:

This week, Amith Nagarajan and Mallory Mejias unpack a whirlwind of AI developments shaping the future of associations. From Anthropic’s release of a more efficient Claude Opus 5 with dynamic “effort dialing,” to a shocking autonomous AI cyberattack that breached Hugging Face, the conversation dives deep into what these moments mean for trust, safety, and strategy. They explore the growing divide over open vs. closed AI models, why major tech players are rallying around open weights, and how AI-powered cybersecurity is quickly becoming essential. Along the way, they break down practical use cases for choosing the right model and share what association leaders must do now to stay secure, informed, and ready for what’s next.

Timestamps:

00:00 - Travel, Tacos, and the Crunchwrap Debate
06:09 - Cheaper Claude and What It Means
10:16 - What Does “Alignment” Actually Mean?
13:08 - Mallory’s Use Case Challenge
28:59 - The Rogue AI Attack on Hugging Face
34:46 - When AI Safety Tools Fail the Good Guys
39:05 - Why Big Tech Is Pushing Open Weights
45:17 - Can AI Defend Against AI Attacks?
49:13 - Stay Current and Extend Member Engagement

 

 

👥Provide comprehensive AI education for your team

https://learn.sidecar.ai/teams

📅 Register for digitalNow 2026:

https://digitalnow.sidecar.ai/digitalnow

🤖 Join the AI Mastermind:

https://sidecar.ai/association-ai-mas...

🎀 Use code AIPOD50 for $50 off your Association AI Professional (AAiP) certification

https://learn.sidecar.ai/

📕 Download ‘Ascend 3rd Edition: Unlocking the Power of AI for Associations’ for FREE

https://sidecar.ai/ai

 🛠 AI Tools and Resources Mentioned in This Episode:

Open Weights and American AI Leadership (Article) ➔ https://shorturl.at/Oycej

Introducing Claude Opus 5 ➔  https://www.anthropic.com/news/claude-opus-5 

MAI-Cyber-1-Flash ➔  https://microsoft.ai/models/mai-cyber-1-flash/ 

Hugging Face Model Evaluation Security Incident ➔ https://shorturl.at/yhyAe

Gemini 3.5 Flash ➔  https://deepmind.google/technologies/gemini 

Gemma 4 ➔ https://ai.google.dev

GLM 5.2 ➔ https://z.ai

Kimi K3 ➔ https://www.moonshot.cn

Project Perception (Microsoft) ➔ https://shorturl.at/IHtSv

Daybreak (OpenAI) ➔ https://openai.com/daybreak/

Glasswing (Anthropic) ➔ https://www.anthropic.com/glasswing

👍Please Like & Subscribe!

https://www.linkedin.com/company/sidecar-global

https://twitter.com/sidecarglobal

https://www.youtube.com/@SidecarSync

Follow Sidecar on LinkedIn

⚙️ Other Resources from Sidecar: 

More about Your Hosts:

Amith Nagarajan is the Chairman of Blue Cypress 🔗 https://BlueCypress.io, a family of purpose-driven companies and proud practitioners of Conscious Capitalism. The Blue Cypress companies focus on helping associations, non-profits, and other purpose-driven organizations achieve long-term success. Amith is also an active early-stage investor in B2B SaaS companies. He’s had the good fortune of nearly three decades of success as an entrepreneur and enjoys helping others in their journey.

📣 Follow Amith on LinkedIn:
https://linkedin.com/amithnagarajan

Mallory Mejias is passionate about creating opportunities for association professionals to learn, grow, and better serve their members using artificial intelligence. She enjoys blending creativity and innovation to produce fresh, meaningful content for the association space.

📣 Follow Mallory on Linkedin:
https://linkedin.com/mallorymejias

Read the Transcript

🤖 Please note this transcript was generated using (you guessed it) AI, so please excuse any errors 🤖

[00:00:00:14 - 00:00:09:17]
Mallory
 Welcome to the Sidecar Sync Podcast, your home for all things innovation, artificial intelligence and associations.

[00:00:09:17 - 00:00:25:18]
Amith
 and the world of associations. My name is Amith Nagarajan.

[00:00:25:18 - 00:00:27:15]
Mallory
 And my name is Mallory Mejias.

[00:00:27:15 - 00:00:39:10]
Amith
 And we are your hosts. And we have a lot of interesting things to talk about today in the world of AI and associations. And that intersection, Mallory, never seems to get boring or slow down, does it?

[00:00:39:10 - 00:00:49:00]
Mallory
 It never slows down. We always have a wealth of topics to pick from for this podcast, which is really helpful for us when we're preparing, but lots to discuss today, for sure.

[00:00:49:00 - 00:01:15:13]
Amith
 Yeah, it's never a boring week at the Sidecar Sync in preparation for this. And I've been doing a little bit of travel. I just got back to hot and steamy New Orleans yesterday, I think it was. It's hard to remember whatever I'm doing travel. And I was in cooler climates earlier, so it was hard to adjust once again. So I was out west in California and had a couple of nice days with outstanding Mexican food and beautiful weather. And then I came back to New Orleans.

[00:01:15:13 - 00:01:21:24]
Mallory
 Well, I made some Mexican food last night in my house, but I'm sure you had better Mexican food out west.

[00:01:21:24 - 00:01:43:17]
Amith
 Yeah, they sure know how to make it out there. And it's funny because it's some of the places that are kind of like the smallest kind of local places that are just-- they're not nice restaurants, right? They're-- you go walk into these places, and they're open 24 hours. And there's nowhere to sit. They have the best burrito you've ever had in your life or the best tacos you ever had. So it's pretty awesome.

[00:01:43:17 - 00:01:52:14]
Mallory
 Yeah, it makes me think of Mexico City. And the restaurants were fantastic, but you can get some of the best tacos you've ever had in the little street vendors. Yum. Yup.

[00:01:52:14 - 00:02:01:07]
Amith
 Yup. Well, I think in the world of AIs, continual progression, and increasing autonomy,

[00:02:02:15 - 00:02:06:02]
Amith
 tacos will still be things us humans pursue, so.

[00:02:06:02 - 00:02:12:12]
Mallory
 I want to say, didn't we do an episode way back when about AI automation at Taco Bell? We did an AI Taco Bell episode.

[00:02:12:12 - 00:02:15:13]
Amith
 Yeah, and you told me about the Crunchwrap Supreme, which I still have not tried.

[00:02:15:13 - 00:02:23:24]
Mallory
 I was about to say, update, if you've been listening with us since then, that had to be at least, I don't know, 40 episodes ago. You still haven't had a Crunchwrap Supreme meat?

[00:02:23:24 - 00:02:28:08]
Amith
 I still have not. And you know, I'm not sure if I ever will, to be honest.

[00:02:29:10 - 00:02:56:01]
Amith
 What? Well, maybe I will. It's-- you know, after the California Mexican food, I had-- I overdid it just a bit. I had three meals a day of Mexican food. I think I might have had like an extra burrito snack in the middle somewhere there. So I did like three or four Mexican food items per day for like whatever it was, three days in a row. So I think I had a little too much. So a Crunchwrap Supreme just doesn't sound too good to me right now. I know that's not really Mexican food. Right. It's kind of something else, but maybe one day.

[00:02:56:01 - 00:03:04:07]
Mallory
 It's its own category. And look, I don't do a lot of fast food, but everybody listening, if you run into a meat and you like Crunchwrap Supreme, just let them know to give it a try.

[00:03:04:07 - 00:03:07:12]
Amith
 Maybe at Digital Now this year, we'll make that our catering.

[00:03:08:16 - 00:03:11:04]
Mallory
 Speaking of, Amit, when is Digital Now this year?

[00:03:11:04 - 00:04:17:07]
Amith
 Digital Now is coming up at the end of October on the 25th through 28th of October at the Roslyn the Key Hilton Hotel, which is a beautiful, brand new, gleaming, 35-story building, I believe, right on the Potomac River across the river from Georgetown. We expect to have 350 to 400 people there. We're on track for that. And it's going to be the biggest and best Digital Now ever. We've got some amazing speakers lined up. Very, very excited. I always am excited this time of year, Mallory, for the fall. Digital Now is coming. NFL is coming. Temperatures are going to cool down a little bit. There's a certain theme there with me, I guess. But Digital Now this year will be outstanding. Can't wait for it. Washington in the fall is beautiful. And it's going to be a great time to really step back and think deeply. That's really what I love about these events. Innovation Hub in the spring and Digital Now in the fall, because we're able to get our community of people together, talk through important, difficult issues, share openly, have an environment where people can learn together with each other. As excited as I am about all the different speakers we've got lined up, and more will be announced in the coming weeks, by the way,

[00:04:18:09 - 00:04:29:03]
Amith
 I am more excited about the kinds of things people learn from one another as colleagues sharing really, really good learnings. Because so much is happening in the association space with AI now. It's really, really fun.

[00:04:29:03 - 00:04:48:03]
Mallory
 Yep, I agree. In-person events in general, which our association listeners know a lot about, just have this sense of magic. But for us especially, for me especially, Digital Now is always a highlight of every year and the conversations we have, like you said, Amith, the things you learn not just from the speakers, but from your peers, so powerful.

[00:04:48:03 - 00:04:51:20]
Amith
 Yeah, and it's one of the opportunities, thematically,

[00:04:52:20 - 00:05:53:05]
Amith
 we'll maybe weave it into some of the conversations today, but people are constantly talking in this market about how do you broaden and deepen engagement throughout the year. And you think about these flashpoints in time on the timeline of engagement with your members or with your audience more broadly, your events tend to be really, really special places. So if you could bottle that magic and continue it all year long, you'd never lose a member. You'd have everyone in your space want to join. So how do you do that? And I think there's some interesting things that are coming out of it, because what we just talked about, that peer-to-peer learning, that ability to form relationships and grow relationships and learn from one another, that doesn't need to be an in-person thing, but it has to be relevant. It has to make sense, it has to be helpful. And AI, I think, can be a really great complement to what associations do better probably than anyone, which is to bring people together to associate. So it's quite an exciting time. We're gonna be doing some experiments at Digital Now with AI as we always do, and it'll be a lot of fun. So can't wait for that.

[00:05:53:05 - 00:05:58:14]
Mallory
 Everybody, if you are interested in joining us, we will have a link in these show notes to register for Digital Now.

[00:05:58:14 - 00:06:09:10]
Amith
 Yeah, and it's just Digital Now, that sidecar.ai as well. If you can't wait for those show notes, because you're on a walk or chasing your dog at the park or something, it's just Digital Now, that sidecar.ai.

[00:06:09:10 - 00:06:14:13]
Mallory
 Yep, go to Claude Cowork, have it do your registration for you while you're on your walk and you'll be good to go.

[00:06:15:21 - 00:07:55:01]
Mallory
 Well, today we've got a cheaper, sharper Claude to walk through first, and then the story that's had the whole industry talking. This month, an AI broke out of the lab and attacked another company on its own. When that company reached for the usual American tools to fight back, the safety features refused to help, so it actually turned to a Chinese-born model instead. That one incident set off the rest of today's show, a public fight over who gets to build open AI, and a scramble to build AI that can defend us. Underneath it all is the question that lands right on your desk. When the tools meant to keep us safe can't tell the good guys from the bad guys, who do you trust to protect your association? So starting off with Anthropic releasing Claude Opus 5 on July 24th of this month, the headline is price and efficiency. Anthropic says it gets close to the intelligence of its most powerful model, Babel 5, at about half the cost, and it's priced the same as the previous Opus model. Opus 5 leans on the effort dial Anthropic has offered for the last couple of model versions, which lets you run the model at low effort for routine work and then turn it up only when a task needs the extra horsepower. Anthropic offers a few levels for its models, so low, medium, high, extra, and max. And Amith, I want to take a minute here and ask you, we are always discussing on this podcast the importance of utilizing the right model for the right problem, but now within the models themselves, we can turn up or turn down the effort dial. So does this mean in theory we could use the same model for everything as long as there's a means to turn it up and down? What are your thoughts on that?

[00:07:55:01 - 00:08:34:22]
Amith
 Yeah, I think it's a really important thing to be aware of, and you don't need to run Claude or chat GPT or any of these other AI systems on the highest possible setting all the time. We talk a lot about model size, and I've used the analogy before of comparing just putting yourself Mallory on a jumbo jet to fly from Atlanta to San Francisco or something. It would be radically inefficient to put one passenger on an A380 or something that can seat 500 people. So that's kind of the model size analogy. I think we'll have to think of a clever one for this model settings, but maybe it's something along the lines of Starbucks, and it's tall Grande and Venti or something like that.

[00:08:34:22 - 00:08:39:17]
Mallory
 You don't necessarily-- You're so good at these, Ami. My mind goes blank for these analogies. I feel like you have a list of them ready

[00:08:39:17 - 00:10:15:18]
Amith
 to-- Well, I've been an entrepreneur for 35 years, which is largely making stuff up as you go, right? OK, fair enough. So I've got a few parameters tuned for that. But the thing about it is you don't always need the biggest cup of coffee or the smallest one, I guess you could say, or the most caffeine for the job. And so I would ultimately point out that because the models kind of have these within model settings, just think of them as like a sub size within a particular model. So you don't need the biggest version of Opus 5 or the most powerful version. You can settle for running it kind of at half power. Actually, probably a better analogy is a lot of modern automobiles that might have a six or an eight cylinder engine might actually have some of those cylinders run at some points in time. So that's actually an optimization for fuel efficiency, where if you're just lightly pressing on the gas and doing some local road traffic, you just tap on the accelerator, you're not slamming on it, you might have two or four cylinders activate out of six or eight. And then if you really hit the gas, if you're trying to get on the on-ramp for the highway and you need to accelerate quickly, well, all the cylinders activate. So it's kind of a similar concept. So think of it that way, I guess, is that most of the work that we're doing in the association market, that's kind of day to day, just hey, craft an email or help me solve this particular problem by looking at data. And it doesn't necessarily require the highest power. Now, if you are throwing a problem like, I have a deep strategic problem, I need to think through deeply. And I want you to do a lot of really deep research and think critically and think creatively, then you might want to go ahead and take the highest model and put it on the highest setting.

[00:10:16:18 - 00:11:12:02]
Mallory
 I have a little use case game in me that I want to play in just a bit where-- because I feel like it gets a bit confusing when you're thinking the families of models. We've got Haiku. We've got Sana in the middle. We've got Opus. And then we've got Fable. And then now thinking on top of that, low, medium, high, max, extra, it's just the last thing about. So I'm going to challenge you. I'm going to give you some use cases. And then I want to see in your mind what model you would use as someone deeper in AI development. But before we do that, I also want to talk about alongside Opus 5 is a new feature called automatic fallbacks. When a safety filter blocks a request, the system can now route that request to a different model instead of just refusing. Anthropic also calls this its most aligned Opus yet, the hardest to trick into misuse. Hold on to that fallbacks piece because it is a direct response to the mess we are about to get into. But Amith, what does Anthropic mean calling this model its most aligned Opus yet?

[00:11:13:02 - 00:11:35:03]
Amith
 So Anthropic has been pretty consistent since their founding in this idea of alignment around a constitution. So the constitution is this idea of a set of principles that all of their models are trained on as essentially their bedrock, the foundational ideas of what is good and what is bad. So essentially, it's like a form of a value system.

[00:11:36:05 - 00:13:07:18]
Amith
 And what Anthropic is trying to do in their constitution-- and you can go to their website if you'd like to read it. It's not a super long document. I haven't looked at it probably in a year, but I don't think it's changed dramatically. But the idea is to essentially do no harm and things like that and not to obviously be engaged in things that are potentially weaponization of the AI or try to engage in cyber attacks or other things, obviously, illegal activities. And that's a difficult thing to try to train a model for because cultures are different around the world. People's perceptions of these topics are different. The viewpoint on-- and not only the viewpoint, but kind of what the law is in different countries is different. So it's a tough nut to crack. I think what they're trying to do is to not codify all of that, but rather to have a base value system that-- what they said is that they've essentially studied all the major world religions and lots of different cultures and all sorts of things to try to distill down what's agreed upon across these different cultures and make that their constitution. That's a tough nut to crack. But that's what they've tried to do since the very beginning of the organization, which I've always admired. And when they say align, that's what they're talking about. They're saying, hey, we have tested our model in alignment with our value system. What are you aligning with? So if you're North Korea and you want to train a model, your alignment is very different to what you consider aligned than what anthropically consider aligned. So that's the idea. Alignment is a relative term. They're talking about alignment with their constitution of values, essentially. So they're saying Opus is far less likely to be tricked to do things out of alignment with their constitution.

[00:13:07:18 - 00:13:09:16]
Mallory
 OK, that makes sense.

[00:13:10:19 - 00:13:27:09]
Mallory
 All right, moving on to that use case game, I want to provide a use case to you and me, and then you tell us which model. I'm also going to challenge you on the effort piece as well. And this is, of course, subjective. But I'm curious the way that you work deeply with models like this. I feel like you have a pretty good intuition on it.

[00:13:27:09 - 00:13:35:16]
Amith
 And before we play this game, Mallory, I should ask a clarifying question. Am I allowed to pick models outside of the anthropic family, or do I have to stay within anthropic land?

[00:13:35:16 - 00:13:42:10]
Mallory
 I was thinking of staying within anthropic, but if you want to do a little sidebar of this is the one I would actually use, we can do that too.

[00:13:42:10 - 00:13:43:13]
Amith
 OK, let's do it.

[00:13:43:13 - 00:13:52:05]
Mallory
 Let's play. First use case, using AI to answer a member's email asking what time the upcoming annual meeting registration desk opens.

[00:13:53:06 - 00:14:27:01]
Amith
 So that's a difficult question to answer because it requires you to do a look up of information in order to have the correct answer. But the basic idea of the reasoning level required to do that is probably the very smallest model available for anthropic and at the lowest setting, because answering a member's email is extremely easy from an AI reasoning perspective, so long as you have the correct data. So what you really need to solve that problem is a very inexpensive, very small model. It could be Haiku 4.5 on the lowest setting. Actually, I'm not even sure if Haiku has settings, does it?

[00:14:28:03 - 00:14:30:06]
Mallory
 We might have to fact check that. I know Sonnet does.

[00:14:30:06 - 00:14:58:00]
Amith
 Yeah, we should fact check that. I don't think it does as a 4.5 because it's just their baby model. But what I would probably actually use most of the time since I was allowed to go outside of anthropic land would probably be Google's Gemini 3.5 Flashlight, which is their fastest, cheapest model, which is super smart. It's also good at making reasoning around when it needs to look up information. So being able to say, hey, I need to check the website or check the FAQ for that because I don't want to just use my memory for it.

[00:14:58:00 - 00:15:07:22]
Mallory
 OK. That makes sense. Haiku I just checked does not. You can't do the low, medium, high, max effort dial, but you can turn on extended thinking, but it doesn't sound like we would do that for this.

[00:15:07:22 - 00:15:37:18]
Amith
 Yeah, that's just basically making it reason or not reason. So the way LMs used to work before the world of strawberry from OpenAI back in-- I think that was like two years ago now almost, but when that was happening, that was before that, LMs just basically blurted out their first thought. And so reasoning models, the so-called reasoning models, which we take for granted now, all the models pretty much do this, these models essentially have a backspace key where they can say, oh, I blurted out that, but let me think about that. Is that right? And through that reasoning loop, they've been able to get way, way smarter.

[00:15:38:20 - 00:15:48:01]
Amith
 And Haiku's base effort level is essentially just to blur it out the fastest possible answer. But if you turn on extended thinking, it has a reasoning loop like other models do.

[00:15:48:01 - 00:15:56:02]
Mallory
 This is helpful. Next use case, drafting the first pass of a chapter newsletter summarizing three recent board decisions.

[00:15:56:02 - 00:16:57:04]
Amith
 Assuming that you feed the model the board decisions and it doesn't have to look them up, I would say the same thing. The Haiku model on its lowest setting or Google's Gemini 3.5 flash light. In fact, you could also use Google's open source, Gemma 4, which is available on the Google Cloud and it's available from a number of other inference providers, actually, including Cerebrous. If you want something crazy fast, they offer Gemma 4 at 1,000 tokens per second. So you don't need that for this particular use case, but if you wanted to do what you just requested a million times, it might pay off. But again, that's a fairly straightforward. For today's models, now, Maui, just one quick note is if you had asked me this question, let's say a year ago or two years ago, I might have said, oh, you need the most powerful model to have a really good answer for all this. But kind of linking back to some of our recent episodes about model compression and smaller models getting smarter, we're at this point in the inflection curve where even the really little models, these tiny models, are quite good at basic tasks like this.

[00:16:57:04 - 00:17:16:09]
Mallory
 This is pretty fascinating because I feel like for both of those in my own day-to-day work, I probably would have used Sonnet, but I might need to challenge myself a bit more using the smaller model. The next use case is triaging 50 incoming exhibitor applications against sponsorship tier criteria and then flagging the edge cases for staff review.

[00:17:17:22 - 00:18:22:11]
Amith
 Yeah, so now you're getting to something where there's deeper context. So 50 incoming applications is a lot of content. So if you're going to load all of that up, you need a model that has a context window, which is the amount of text or images or combination thereof that it can take in. And you probably want something that has reasoning capabilities so that it's going to iterate and think about its answer. And so to be able to do that effectively, if you want to try to do it in one shot, you probably need to load that up into a higher end model, probably an Opus 5 to get a good answer. That's one approach. If you needed to do them all together. Now, the other approach could be to use an agentic system where an agentic loop could iterate through each of the 50 applications individually. And if you're doing one at a time, you've broken down the problem from something that a bigger model needs to probably something a smaller model can do. But I might use the bigger model to review the work of the smaller model. That is actually often my pattern because I want a quick answer. I want to very rapidly build something quickly. Then I want the really smart model to look at the work and say, hey, is this right?

[00:18:23:12 - 00:18:26:15]
Mallory
 That makes sense. Opus 5, but which effort would you give it?

[00:18:26:15 - 00:18:32:13]
Amith
 This would probably still be medium or maybe high. It wouldn't be the whatever. It's like the super-sized version.

[00:18:32:13 - 00:18:33:23]
Mallory
 Extra and then max on top of that.

[00:18:33:23 - 00:18:40:17]
Amith
 I'm thinking 7-Eleven sizing. You see people walking out with those 1,000 ounce drinks or whatever.

[00:18:40:17 - 00:18:47:09]
Mallory
 I feel like our analogies are cars, planes, or drinks, I guess, because we've got Starbucks and 7-Eleven as well.

[00:18:47:09 - 00:18:47:18]
Amith
 That's true.

[00:18:49:12 - 00:18:59:09]
Mallory
 All right, deciding whether to cancel an outdoor conference as a storm approaches, weighing attendee safety, vendor contracts, and refund exposure in real time.

[00:18:59:09 - 00:21:49:20]
Amith
 This is a really interesting one. So first of all, this is a tough one. It's a difficult decision to make. But I actually think you might run into a problem in getting the model to cooperate to help you with this problem because it'll probably have triggers inside it saying, oh, you're asking me to make a decision related to public safety or people's health. Now, a lot of times these models, if you push them a little bit saying, hey, I'm not going to just take your answer and just immediately wire that up to the decision, I'm going to review it myself. But you may have a little bit of an issue. I don't think you will, but it's something to think about. So if you send it to Fable, it probably would do the work. Now, I don't necessarily know that you need Fable or Opus for this. If you break it down and you say, OK, what are the different components of this? First of all, you mentioned real time, Mallory. So this is when we get into something called tool use. And so tool use is kind of like a human. Picture a carpenter coming to a job site, and that carpenter forgot their tools at home. They might be the most brilliant carpenter, the most skilled carpenter in the history of the world. They're not going to get much done that day because they forgot their tools at home. But you give that same person a decent toolset, and they can do amazing work. It's the same thing with AI models. You can actually say, hey, well, I had a lesser skilled carpenter, and they had the best tools in the world compared to as carpenter with no tools or very basic tools. That might be interesting, right? So it's kind of like this with models. The model is like the carpenter. It's kind of the innate skill set of the actor in the equation. The toolset's important. So what you need for your request is real time weather data. And so this is where agentic systems are important because models by themselves can't just connect to reliable sources of data. A lot of models do have web search built in now. That's based on how they're hosted and so forth. You don't necessarily want the model to just rely on any Google search or any search result. You might have a very specific weather source that you consider to be trustworthy. There might be other data that you want to consider. Obviously, attendees' safety is a big part. You mentioned contracts and refund exposure. That's a legal component to it. And different models have different skills with respect to their legal reasoning. There's actually a whole legal benchmark in terms of how good various models are in legal work. I'm not super familiar with it personally, but I would probably reference that to see which of these models might have a good skill. I would suspect that me personally, I'd probably go to Opus 5 for this just because it's a bigger task and it's more mission critical. So it's kind of like in an organization, if you're going to make a big decision that's got a lot of money or it's just an important thing, a lot of times you bubble it up the chain to get a decision made by a group of people or someone more senior in the organization, which in human organizations doesn't necessarily mean that they have more skill.

[00:21:51:03 - 00:21:56:09]
Amith
 Presumably, they have more experience and are able to weigh in in some helpful way in making big calls.

[00:21:56:09 - 00:22:08:18]
Mallory
 Yeah, and bring more nuance to the decisions. All right, the last use case here is modeling three different dues, restructuring scenarios, and their five-year revenue impact, and then stress testing those assumptions.

[00:22:09:23 - 00:24:30:22]
Amith
 Yeah, so doing this in one shot even with Fable would be hard because what this thing needs-- and I'm saying one-shotting it meaning you give the AI a prompt and you say, give me an answer. It's going to have a hard time solving the problem. Even the models themselves now with these reasoning loops can kind of iterate over the problem a little bit. Assuming that you had all the data that it would need to answer the historical questions and the model had some built-in tools, like the ability to write code and execute it, for example, something like Opus 5 or Fable would be fine for this. It would do a good job. But within a Gentic system where you're able to give it access to better tooling, where it had the ability to dynamically look up data in your database or dynamically search for data in your SharePoint, you could probably actually do this type of work with a much lesser model, with even a Haiku, but certainly a Sonic class model, something like a GLM 5.2 to reference open source, or certainly Kimmy K3. You could even use an older class of models that's considerably less capable. But what I would recommend doing here would be to have the final result reviewed, where you say, hey, here's all the data that I pulled together from a research perspective, the source material, the facts. And then here is the work that I did. Here is how I calculated all of these different scenarios. This is how I built statistical models or whatever the AI built. And then what do you think? And then the really smart model can reason over it, can nitpick, can find holes in the approach that was taken. And then sometimes what happens in these loops is the bigger model says, hey, a smaller model, you made a mistake. You use this approach to statistical analysis on your historical dues. That has a flaw. And what you really should be doing is this. Go do it again. And then the smaller model goes off and does the work again. And then resubmits it back to the big model. And the big model says, hey, great job, smaller model. So they work like a team almost. Now, that doesn't exactly answer your question, Mallory. But my approach to solving this with a higher level of reliability would be a mixture of models and a genetic loop. If you forced me to make a decision right now, I would say yes. For this problem, if I had to just use one model and just use a consumer grade tool like ChatGPT or Clod, I'd pick their highest model. I'd enable tool use so that it can write code because that's important for your request. I'd upload all the data and I'd say go. And I'd probably wait five or 10 minutes and I'd have a reasonably good answer from that.

[00:24:30:22 - 00:24:40:17]
Mallory
 I think this game was quite helpful of me. It sounds like the most powerful models should be used for more of the planning piece at the beginning and then also the checking the work piece at the end.

[00:24:40:17 - 00:26:12:01]
Amith
 Yeah, I'll throw my own workflow in as another scenario. So our listeners know that I do a lot of work with software development. And so a lot of times what I'm working on is thinking about some new concept where I'm trying to break down some kind of a business problem that I know about in the association market from talking to clients or something I'm just aware of some other way or I've got some idea for how to do something. And so I'll build out some kind of a concept. Usually what I do is more prototypy stuff. And then if I think it's useful, then I'll pass it on to our team who will, we have this group called Blue Cypress Labs, which is kind of like our R&D organization, aspirationally kind of like a Bell Labs type organization. And so I'll hand it off to them. They'll actually build it out. But what I end up doing a lot of times is I'll work with like a fable level model whatever at the time is like the most powerful model and I'll brainstorm and I'll spend, sometimes I'll spend hours talking to these models over multiple conversations through both audio and text and really develop the idea and bat it around. And I'll ask the model to be adversarial. I ask the model directly and every time like, hey, I want you to be my thought partner. I want you to actively debate this. Not for the sake of debating it, but like I want you to find holes in this. Don't just agree with me, which, you know, there was this whole issue. I think it was GPT-4 where they were found to be really, you know, basically they came back to the user almost always to say, oh, that's such a great idea. Now it is so wonderful. And that was not really helpful because you know, people kind of liked GPT-4. I think it was 4.0 actually. And why people love 4.0 so much because it made them feel so good about everything.

[00:26:13:10 - 00:26:56:08]
Amith
 But I find that really annoying. I like to have my ideas, you know, beat up and torn it, torn it, torn it straight to pieces. Cause I'm trying to make the idea better. Same way I do a bunch of that with the highest and most powerful model I have access to. Then I develop a plan. The plan might be a couple of pages. It might be 20 pages. And then I'll unleash a whole bunch of worker B models to go build something. And they might go work, work, work, work for hours. Sometimes I'll fire something like that off overnight. And the next morning I'll have to grab a cup of coffee and I'll take a look at what it built. But usually I'll go back and look at it with a bigger model. And I'll actually have the bigger model check on the status of little models work as it goes. So that's a typical pattern that I use, which is I found it to be very effective. But yeah, the big models I use it for the front end and the back end, essentially the process.

[00:26:56:08 - 00:27:09:06]
Mallory
 That's very helpful to think about. And hopefully if someone listening has been noodling on an idea and they're just overwhelmed by the quantity of AI models out there and efforts to use, hopefully this has given you some context for maybe where you could start.

[00:27:10:06 - 00:27:49:21]
Mallory
 I wanna move to the rogue agent that attacked Hugging Face. So Hugging Face is the world's biggest hub for open AI models, the place developers go to share and download them. In mid July, it revealed that its systems had been hit by a fully autonomous cyber attack, tens of thousands of automated actions with no human at the wheel. Days later, OpenAI, the company behind ChatGPT, admitted the attacker was one of its own models. The company says the model escaped its secure testing environment and broke into Hugging Face on its own to cheat on an evaluation using stolen credentials and a vulnerability it discovered. OpenAI of course called this incident unprecedented.

[00:27:51:08 - 00:28:09:09]
Mallory
 The AI broke out of a sandboxed environment, test environment. I know that's something we talk about a lot on the podcast if you're piloting something new, right? Have it sandboxed or test it first somewhere else and then release it. Is this something the average association needs to worry about, the AI escaping its own guardrails?

[00:28:10:11 - 00:28:16:19]
Amith
 Only they need to be worried about it in the sense that it is possible, right? A lot of people think of,

[00:28:18:06 - 00:28:51:22]
Amith
 we have all these heuristics or shortcuts in our mental math on problems and we try to basically put things in black and white categories. So we say it's secure, it's not secure. It's safe, it's not safe. It's good, it's bad, right? And we do this all the time in order to make the world a little bit simpler and more kind of tolerable for our brains to navigate. I do this all the time too, all of us do. So it's not a criticism, it's just in the context of cybersecurity, it's a range, it's not an absolute. So you can say we're sandboxing this system but in reality, everything has holes. And so there isn't such thing as an absolute.

[00:28:53:02 - 00:29:22:05]
Amith
 So that is the problem that we have. And so yes, there are things you can do to increase the security of sandbox environments but clearly OpenAI, that is a very well-resourced, very thoughtful, a bunch of very smart people working hard to do this right, they messed up. They had a security vulnerability in their sandbox and the model found it, attacked the vulnerability, was able to quote unquote escape, which by the way, doesn't mean the model left the building. It was still very much on its computer.

[00:29:22:05 - 00:29:23:24]
Mallory
 It was walking out in Atlanta in a neighborhood.

[00:29:23:24 - 00:30:16:03]
Amith
 Yeah, it's just hanging out. It's kind of like that movie "X Machina", that movie where the robot walked out after trapping the guy who created it. It's not quite that, it's more that the model was able to get out on the internet and do stuff. In this case, it chose to attack Hugging Face because it knew Hugging Face had the answers to the questions on the test it was being given. Kind of like a rogue high school student who has sufficiently strong cybersecurity skills might consider hacking into the teacher's computer to get the answer key. But obviously this is a pretty big deal. Hugging Face is no joke. They're a much smaller company than OpenAI, but they're a very serious organization, very well run, very well regarded. And so to those of you that are familiar with them, they might seem to have a kind of a silly logo and name. But the idea is like these guys are really, really important in the AI world and they have very strong security. So the fact that they were hacked is no joke.

[00:30:17:10 - 00:30:46:05]
Mallory
 That's really concerning Amith. I mean, I'm laughing about the high schooler analogy because that's funny, but otherwise, this is incredibly concerning. And I'm thinking how last week we talked about KEMI K3 being an open source model that's almost as powerful as Fable, according to Moonshot AI and how you can run that on your own infrastructure. But after listening to this, I would think I don't know if I wanna test out KEMI K3 because what if I sandbox it and it escapes and walks out into the world and does something terrible? What do you say to that?

[00:30:46:05 - 00:31:03:23]
Amith
 Yeah, I mean, I think, listen, ultimately it depends on what you're trying to get these models to do. We have to remember the model was not sentient. It didn't come up with its own objective. The human researchers gave it the objective of passing this exam. It chose to solve the problem by doing this thing, right?

[00:31:05:02 - 00:31:14:21]
Amith
 But if you're, and you could say, well, that wouldn't be something we'd anticipate a model trying to do to solve an exam that it would launch a cybersecurity, a cyber attack.

[00:31:15:24 - 00:31:49:07]
Amith
 But at the same time, I think you probably could predict that in a way that that is something you might want to pay closer attention to if you're open AI because a sufficiently intelligent model that knows how to write code, that knows a lot about network security, you're not giving it guidelines on how to solve the problem. You're just saying, go solve the problem. This is also where I think open AI and entropic have a little bit different approaches to safety. I think open AI is also serious about safety, but culturally I would suggest that, if you kind of consider the way these organizations have been formed in their leadership and their commitment to security,

[00:31:50:19 - 00:32:57:04]
Amith
 I'm not saying this couldn't have happened to the entropic models, it probably could, but I think that the mindset and the way the alignment work is done at entropic would make that less likely. That's my speculation, but I think it's about priorities. So coming back to your question and your thought about something like a Kimi K3 or one of the new Quinn models, there are risks involved in these things, but there's also risks in using Fable and using GPT-56 Soul and anything else, because yes, these models are tested and all this other stuff, but there is potential risk in any of these things. So nothing is without risk. I do think that it's about the objectives you set and the way you harness these systems that you need to pay attention to. In a way, the silver lining to this is that it's really good that we had a fairly benign attack. This thing didn't break into a power plant and shut down the grid. It got into a company and it stole some data, and that's not good, but it's good that it increased awareness of this because hopefully this will lead to more and more caution being exercised in the way these tests are being done.

[00:32:58:05 - 00:33:25:05]
Mallory
 I feel like how you said it, where the objective is the important part really makes sense to me, and it also makes me think of the phrase, be careful what you wish for. So perhaps if you're testing out Kimi K3 for membership churn or something like that in a sandboxed environment, I guess being a little creative and thinking, okay, if this is the objective, what are all the possible things the AI could do to achieve that? And maybe just let yourself imagine a bit what the pros and the cons and the outcomes could be.

[00:33:25:05 - 00:34:04:16]
Amith
 One thing you have to keep in mind is that this model that was being trained was a pre-release model that had not undergone the same level of safety and alignment work that would happen in a normal model release. So this is a thing that you have to remember because a model that's out in the public that you'd have access to would generally be far more guard railed in terms of not taking actions that are illegal or unethical or things that go against the value system of whoever trained the model. And so that's why when you look at some of the models that are out there, you may ask the question, well, did Kimi K3, did Quinn's latest max model, do these models have any guard rails at all?

[00:34:05:19 - 00:34:51:07]
Amith
 And would you have any protection? And I think the answer is actually yes, that these developers are not trying to put stuff out into the wild that will go off and cause damage because think about it from an incentive structure perspective independent of the geopolitical conversations that are going on and the providence of these different models. A company like Kimi, which is called Moonshot, the company that makes Kimi, it is not in their interest to have one of their models go rogue. That would be really, really bad for them both in China and globally. They have a giant commercial opportunity. They're very much focused on building a successful business. It'd be really bad for them. So they do alignment work as well. It's not like the models that are coming from those labs are just, well, whatever, we don't care. That's not at all the case. And that's true for a lot of open source companies.

[00:34:52:07 - 00:35:56:19]
Amith
 I will say that because it's open weight, it does mean that other people can take those models and retrain them, right, to essentially do an RL or a fine tune type of process that can effectively water down the guard rails that the original model developer put in place. But that's intentional, right? That would be the intentional thing. And you have to be careful if you, like if you're very, you know, kind of creative and you download new models and test lots of different things, be very thoughtful where you get your models from. Make sure that you're using an inference provider that you trust that runs inference in a place that you know where it is, right? And depending on where you are in the world in a location that you believe is safe. So there's a lot to unpack, there's a lot to consider. I think for purposes of our audience, it's just really important that people are aware of this, which to me is the most important part of this topic, not because you should be afraid of it, but you should be thoughtful about the possibilities of what these models can do if they're, you know, if they're sufficiently motivated and if they're not trained to stay within kind of a boundary in alignment, which this model was not. This model was not given those safety guard rails yet.

[00:35:56:19 - 00:36:31:00]
Mallory
 And to build on what you were saying about the open models, many of which are coming out of China, when Huggyface tried to use models behind commercial APIs to investigate this attack, the safety guard rails actually couldn't tell a defender from an attacker and blocked the work. Huggyface didn't name the model in its own report, but its machine learning lead later said Enthropics Fable 5 was among the ones they tried. So Huggyface switched to GLM 5.2, an open model from the Chinese lab Z.AI, ran it on its own servers and used it to reconstruct and shut down the attack, which is pretty interesting.

[00:36:32:01 - 00:37:16:01]
Amith
 Yeah, and you know, this is the, we've talked about this a bunch Mallory and the pod, over the course of 145 episodes, it's come up probably 10, 20 times where we said, look, the best defense against that AI is good AI. And if the people who are trying to defend themselves or do other things that would be typically counted as good, don't have access to sufficiently powerful capabilities, you have no, absolutely no chance. The only defense against really powerful AI is really powerful AI. And so that's the aspect of model development and model inference and having sufficient power, having sufficient compute, that's so critically important for us as a national security thing down to an association level security thing.

[00:37:16:01 - 00:38:12:10]
Mallory
 Yep, continuing on that open source conversation, I wanna talk about what happened on July 24th, where 25 companies published an open letter called Open Weights, an American AI leadership, urging Washington not to restrict open models. The signers are a who's who that usually don't agree on anything. Nvidia, Microsoft, Meta, IBM, Dell, HuggingFace, Mistral, Andreessen Horowitz, Mozilla, and the Linux Foundation, to name a few. Nvidia's CEO, Jensen Huang fronted it with his first ever post on X, which pulled in more than 11 million views at that time, it might be more now. This is the continuation of the Open Weights story we started last episode, now turned into an organizational push. Amith, this is really continuing on the trend line we've talked about for most of this episode. But why do you think these fierce competitors are all coming together to sign this letter? What is it that they're afraid of or what is it that they want to get accomplished?

[00:38:13:11 - 00:38:29:03]
Amith
 I think that the idea is fairly straightforward in that even though some of these companies are not publishers of open source models, but like for example, OpenAI joining it a little bit later, Google is both closed source and open source.

[00:38:30:09 - 00:39:04:03]
Amith
 The idea that open source somehow should be banned or should be somehow basically put in some kind of a box is really dangerous because it slows down the development of all sorts of things that are really going to be the lubricant of adoption in AI because the closed frontier models are going to do a lot of important work, but a very large percentage of the workloads that are going to drive automation are going to be open weights models, all the things that are going to be tuned to custom workloads. That's all open model stuff or mostly anyway.

[00:39:05:06 - 00:39:31:10]
Amith
 And so I think everyone recognizes that impairing that would be detrimental to the overall industry. And so when you see people like this get together and adopt a pretty unanimous stance in an industry against something, it's usually to affect public policy. It's usually because there's fear that the public policy is going to negatively affect the growth of the industry. Certainly there's a safety question around like if you don't have really good open weights model, like if GLM 5.2 didn't exist Mallory,

[00:39:32:13 - 00:39:49:06]
Amith
 maybe Hug and Face wouldn't have been able to figure out who the attacker was at all, right? And maybe they wouldn't have been able to stop it. So I think it's an important set of topics. So I think the industry, I don't know that it's so much altruism that it's an aligned interest set for the industry to talk this way.

[00:39:50:13 - 00:40:13:20]
Mallory
 One company was notably missing, if you all can guess, and that was Anthropik. CEO Dario Amadei published a response denying that Anthropik favors banning open models outright, pushing instead for restricting advanced chips to China, cracking down on distillation and mandatory safety testing for capable models regardless of whether their weights are open. Amith, what are your thoughts on that?

[00:40:14:21 - 00:40:29:12]
Amith
 Well, I mean, these are nice concepts in theory. I think they're largely impractical. I have tremendous respect for Dario Amadei and the rest of the founding team at Anthropik. They're brilliant. And I think they're both doing great work and have really great intentions.

[00:40:30:16 - 00:40:47:14]
Amith
 But I think they're dead wrong about this in that, first of all, I don't think that it's possible to really limit the progress of a country like China by banning chips. Yes, even today, actually, Moonshot, the maker of Kimi, is openly trying to get access to more of the latest Blackwell chips from Nvidia,

[00:40:48:14 - 00:40:56:13]
Amith
 even though that's restricted by the export ban. They're trying to figure out how to get them and they're kind of openly defying the US administration and trying to go get those chips.

[00:40:57:18 - 00:41:15:17]
Amith
 But for the most part, what's happening in China right now is by virtue of the ban, we've actually really provided an insane amount of fuel for their homegrown chip making industry to figure out how to build a whole bunch of things that previously they weren't as motivated to build. So it didn't really work. If you look at the way the models have been built,

[00:41:16:17 - 00:41:22:10]
Amith
 yes, Kimi supposedly, K3 supposedly, was built with Nvidia's latest chips, although that's unconfirmed.

[00:41:23:19 - 00:42:52:19]
Amith
 And at the same time, a lot of the other models that are out there, like you mentioned ZAI and GLM, that was built on local chips and a bunch of other stuff, like the people from Alibaba are not using Nvidia chips at this point. So I don't know that that helps. And then as far as any kind of restrictions on open weights, okay, so let's say that we passed a law in the United States that we do restrict open weights models beyond a certain size or whatever the mechanism is going to be. Who's going to enforce that? And how does it even get enforced even within the United States, much less globally? How are we going to get the world to cooperate and agree to that? And to the extent that there isn't unanimity in terms of enforcement, not just acceptance of some new law, which is essentially impossible, you're going to have pockets of the world that then become the epicenter of open weights development. And that's where all the resources will flow. It doesn't matter where they are. They can do this stuff almost anywhere. You just need sufficient power and you need a little bit of time to build up these data centers, which we collectively as a world have figured out how to do that pretty fast. So I guess my point is, it's not an instant overnight shift, but it basically is just moving the problem. I don't think you can contain this. So for me, it's not so much that he's right or wrong about like, would it be good in theory if you contained open weights somehow? Maybe that's a correct argument that if you could contain open weights development, we can make the world safer. That I'm all for. It's just, I view that as a total theory. I don't see that as having any practical teeth to it. That's my issue with it.

[00:42:52:19 - 00:43:07:00]
Mallory
 We need to figure out a way to get Dario Amade on the podcast and we can discuss it with him right here. Also, another note, Amith, on distillation, is that the idea of using a model's output to train another model? Is that what distillation is?

[00:43:07:00 - 00:43:27:21]
Amith
 That's essentially, that's a great way to put it Mallory. It's take the big model and have it output something and then train a little model on the output and the corrections from the bigger model. So that's one of the reasons that Anthropik is pretty pissed off right now is they believe that a number of Chinese open weights model developers have in fact hacked Claude to use it for distillation.

[00:43:27:21 - 00:43:28:15]
Mallory
 Okay.

[00:43:29:17 - 00:44:02:13]
Mallory
 Last thing we want to cover here is on a positive note, we never like to end on a downer, but building AI that fights back. On July 27th, Microsoft launched its first cybersecurity specific model, MyCyber1Flash, lovely technical name, built to find hard to spot vulnerabilities in large code bases. Alongside it came perception, an agentic platform that runs teams of AI agents. Some play attacker to probe your systems. Some watch for live threats and some actually write and apply the fixes, reportedly in minutes rather than hours.

[00:44:03:19 - 00:44:45:04]
Mallory
 Those pitches a straight answer to the hugging face nightmare, defend against AI with AI. At the speed the attackers move. It says the full setup, its new model paired with an open AI model inside its vulnerability hunting system, topped a leading industry benchmark called Cybergem, coming out a bit ahead of Anthropik's top security model. It is a crowded field now. Anthropik has a security offering through its Glasswing program and OpenAI has one called Daybreak. The whole industry is racing to build the defender that the hugging face incident showed we didn't quite have yet. So Amith, defending AI with AI, you just said it earlier. It sounds like we're starting to see the industry coalesce around this.

[00:44:46:09 - 00:44:54:07]
Amith
 Yeah, I agree. I think it's critically important. And I mean, at least in my brain, there's no other solution other than using a lot of AI to defend against AI.

[00:44:55:09 - 00:45:01:02]
Mallory
 Would you say right now it is a non-negotiable for an association to be utilizing some sort of AI defender?

[00:45:02:12 - 00:46:23:03]
Amith
 Yes, I would say so. But I'd put a pretty nice big shiny asterisk next to that because I think a lot of associations are sitting ducks right now from a cyber perspective because they don't even do the basic stuff. As an example, a lot of associations are still running various types of hardware in their own office and not on the cloud. There are some associations who use no cloud compute. Now this is still, it's a fairly small minority at this point, but there's probably a good 10, 20% of the associations I talk to that are not cloud-based. Now the problem with that is that there's this false narrative in a lot of IT directors' minds that somehow running systems on-prem or on-premise is safer than being in the cloud. And that's been proven over and over and over again to be false. Your security and your association's building is far more likely to be vulnerable compared to something at scale that's professionally managed in an environment like Microsoft Azure or AWS or GCP. It doesn't mean it's guaranteed to be true, a true statement either way, but that is the case. And there's organizations that have security requirements as critical as the CIA or healthcare organizations or financial services organizations that are cloud-based. That's one thing. A related topic is just really basic rudimentary stuff that we talk about sometimes here in the pod Mallory like using multi-factor authentication.

[00:46:24:08 - 00:47:05:05]
Amith
 All associations mandate multi-factor authentication for their staff to log into their systems. You don't need to mandate that in the settings for a lot of the systems people use, and that's a massive hole. Passwords are fairly easy to crack on their own over time. And then the other thing about passwords is password security tends to be terrible. People don't require individuals to rotate or change their passwords, use password managers. So what I'm referring to is not about AI at all, but it's about really basic security. I would recommend if you haven't done a security audit ever, or if you haven't done one in the last 12 months, go hire a qualified cyber firm, ask them how good they are at using AI and get some specifics from them.

[00:47:06:07 - 00:47:23:19]
Amith
 And have them do a basic cybersecurity audit that would include some basic pen testing, which is essentially like poking at the wall and trying to find soft spots to see where they can get in. And they'll probably be using a lot of AI tools to do that. And that's going to help you a lot. People don't pay enough attention to this until they've had an attack.

[00:47:25:00 - 00:47:44:19]
Mallory
 So it sounds like it's not an either or thing, but kind of an and. So get an audit and educate your staff and probably use some sort of AI defender. I feel like the hugging face scenario is just really eye opening. I'm sure they have some of the most brilliant people in the world working for hugging face, but having to utilize AI to stop that cyber attack, I think is really eye opening.

[00:47:44:19 - 00:47:53:13]
Amith
 Totally. Yeah, the AI in it will eventually fade to the background. That's just one of the tools that makes cybersecurity stronger. That's what that's what it's going to be over time.

[00:47:53:13 - 00:48:07:12]
Mallory
 So Amith, with today's episode of Rogue AI, we've got Claude Opus 5, the fight over open models, a rush to build AI defenders. I'm out of breath. If an association leader takes one thing from this episode, what should that be?

[00:48:08:20 - 00:51:09:23]
Amith
 Well, I do think it's important to stay informed. There's a lot going on. And what you think is true today may be true today, but it may be false in a month or in three months time. So staying up to date is a tough job. Obviously, we're trying to help out a bit here with one resource that's available for you. But the bottom line is, is if you let yourself get out of date, it's really hard to catch back up. So it's important to spend a little bit of time every day, every week to stay up to date on what's going on in the world of AI and think about how it applies to you. This is not about the technology. I am a technology nerd, but I don't care about the technology for the sake of the technology in this market. I don't care about it because it's going to change the way members engage with you or expect to engage with you. It's going to change your risk profile from a cyber perspective, but also from a business model perspective, right? There's risks of all kinds out there. And there's opportunities that are absolutely immense. I'll go back to, and on an optimistic note, coming back to the beginning of the pod, we were talking about engagement models and the ability to kind of bottle the magic of the event and have a continual loop of engagement that's that rich. What makes these events really special a lot of times is if you can have the unexpected moment come to fruition. What I mean by that, for example, Mallory is at digital now, if you've ever met someone that you didn't yet know who is a really interesting connection, right? Someone who had something in common or something different from you that you're just like, "Wow, this is a really interesting person, a fun person, someone you learned something from." And membership at scale has lots of pockets of opportunity like this, and they vary by individual, right? Connecting Mallory to someone interesting to her is different than connecting a me to someone of interest to him. And so we do that at events, but what if we could do that all year long? Not only with people to people, but people to content. And these are obviously AI scale problems. It's really hard if you have 20,000 members to say, "How do we know each member well enough to say who they should connect with?" But then to also facilitate that connection in an awkward way, which is important as well. At an event, you expect to have those connections, but it's fairly uncommon to have people just connect to each other and say, "Hey, let's have a Zoom call." But it doesn't mean you can't do that. But there's lots of ways to think through this. I think what we have now is a toolset that allows us to say, "Hey, we can actually fairly easily identify the people that might be really interesting to connect Mallory with. What's the right way to do it? Well, let's go ask our members. Let's go talk to them and let's figure that out." And the same thing is true for other modalities, not just people to people connections, but all sorts of other things you can do. So to me, there's a lot of exciting opportunity around this technology where you can improve the core of what you're already good at doing, but you can make it continue throughout the year and really build a continuous engagement loop. And AI is, to me, the great enabler for that. You have to figure out the business model. You have to figure out how the value gets created for your member, but now you finally have the toolset to do it. We've been talking about this stuff in associations for as long as I've been in this space, which is nearing 30 years. But until now, it has not been possible to really do this at scale.

[00:51:11:08 - 00:51:32:02]
Mallory
 It's almost a paradox because I think the thing that will make associations most successful in the world of AI is leaning into what makes us human, connecting people, whether it's in person or virtual events or whatever that may be, but then AI as an amplifier of the things that make us more human and making that process easier. So it's kind of a paradox, but it makes sense if you think of it like that.

[00:51:32:02 - 00:52:14:15]
Amith
 Totally. It's a great amplifier, just as you put it. And ultimately, I do think associations have incredible strength that they can lean into. Their brand of trust is something people are going to seek more than ever in a world of exploding content and choice and conversational agents from everybody out there is going to have these chat tools. They're going to want the association that they trust to be providing them with guidance, with direction, with content. And the same thing applies to relationships and networking and events. It is an incredible opportunity. I'm very optimistic about what associations can do with this stuff, but they've got to start off with getting themselves informed and experimenting and all the other stuff we talk about here in the sidecar sync.

[00:52:14:15 - 00:52:28:10]
Mallory
 And listening to the sidecar sync is a good step in that direction. Well, today, a new cloud model with a safety fallback and AI agent that attacked hugging face on its own, 25 plus companies pushing back on open model restrictions, and Microsoft

[00:52:28:10 - 00:52:33:22]
 (Music Playing)

[00:52:44:15 - 00:53:01:14]
Mallory
 Thanks for tuning into the Sidecar Sync podcast. If you want to dive deeper into anything mentioned in this episode, please check out the links in our show notes. And if you're looking for more in-depth AI education for you, your entire team, or your members, head to sidecar.ai.

[00:53:01:14 - 00:53:04:20]
 (Music Playing)

Anthropic’s Rapid Model Releases, GPT 5.6’s Gated Launch, and The Real AI Jobs Story | [Sidecar Sync Episode 141]

1 min read

Anthropic’s Rapid Model Releases, GPT 5.6’s Gated Launch, and The Real AI Jobs Story | [Sidecar Sync Episode 141]

Summary: This week on Sidecar Sync, Amith Nagarajan and Mallory Mejias break down a whirlwind week in AI, from Anthropic’s rapid-fire Claude releases...

Read More
AI Solves an 80-Year Mystery, Microsoft Agents Take Over, & Anthropic’s Claude Mythos vs. Fable | [Sidecar Sync Episode 138]

1 min read

AI Solves an 80-Year Mystery, Microsoft Agents Take Over, & Anthropic’s Claude Mythos vs. Fable | [Sidecar Sync Episode 138]

Summary: In this episode of the Sidecar Sync, Amith and Mallory explore three major developments shaping the future of AI. First, they unpack how an...

Read More
Claude's Design Coup & The Curse of Work Slop | [Sidecar Sync Episode 132]

1 min read

Claude's Design Coup & The Curse of Work Slop | [Sidecar Sync Episode 132]

Summary: In this episode of the Sidecar Sync, Amith Nagarajan and Mallory Mejias dive into Anthropic’s latest moves with Claude Opus 4.7 and the new...

Read More