Summary:
AI may feel like it arrived overnight, but its foundations have been decades in the making. In part one of a refreshed Foundations of AI series, Amith Nagarajan and Mallory Mejias unpack the essential concepts association leaders need to understand in today’s rapidly evolving AI landscape. They trace the convergence of compute, data, and algorithms that fueled the generative AI boom, challenge assumptions about whether we’ve already reached some definitions of AGI, and explain the difference between training and inference. They also break down large language models, multimodality, open versus closed models, and the weights that form the heart of artificial neural networks. Along the way, Amith makes the case for building flexibility into your AI strategy as models and inference providers continue to change at breakneck speed. Whether you’re new to AI or ready for a fundamentals refresh, this episode provides a practical foundation for understanding what’s happening under the hood—and why it matters for associations.
Timestamps:
00:00 - Revisiting the AI Fundamentals, Two Years On
04:41 - AI Is 70 Years Old, Generative AI Is Not
09:58 - Overhyped Now, Underhyped Later
16:05 - Training Versus Inference, Explained
23:28 - Is an LLM Just Fancy Autocomplete?
26:19 - What Multimodal Actually Means
29:58 - ImageNet, AlexNet, and the Deep Learning Breakthrough
32:03 - Open Weights Versus Closed Models
34:54 - What Model Weights Actually Are
37:52 - Recapping Part One and What's Next
👥Provide comprehensive AI education for your team
https://learn.sidecar.ai/teams
📅 Register for digitalNow 2026:
https://digitalnow.sidecar.ai/digitalnow
🤖 Join the AI Mastermind:
https://sidecar.ai/association-ai-mas...
🎀 Use code AIPOD50 for $50 off your Association AI Professional (AAiP) certification
📕 Download ‘Ascend 3rd Edition: Unlocking the Power of AI for Associations’ for FREE
🛠 AI Tools and Resources Mentioned in This Episode:
Claude ➔ https://claude.ai
Claude Code ➔ https://claude.com/product/claude-code
ChatGPT ➔ https://chatgpt.com
Gemini ➔ https://gemini.google.com
Gemma ➔ https://deepmind.google/models/gemma
Kimi ➔ https://kimi.com
Fireworks AI ➔ https://fireworks.ai
Cerebras ➔ https://www.cerebras.ai
https://www.linkedin.com/company/sidecar-global
https://twitter.com/sidecarglobal
https://www.youtube.com/@SidecarSync
⚙️ Other Resources from Sidecar:
More about Your Hosts:
Amith Nagarajan is the Chairman of Blue Cypress 🔗 https://BlueCypress.io, a family of purpose-driven companies and proud practitioners of Conscious Capitalism. The Blue Cypress companies focus on helping associations, non-profits, and other purpose-driven organizations achieve long-term success. Amith is also an active early-stage investor in B2B SaaS companies. He’s had the good fortune of nearly three decades of success as an entrepreneur and enjoys helping others in their journey.
📣 Follow Amith on LinkedIn:
https://linkedin.com/amithnagarajan
Mallory Mejias is passionate about creating opportunities for association professionals to learn, grow, and better serve their members using artificial intelligence. She enjoys blending creativity and innovation to produce fresh, meaningful content for the association space.
📣 Follow Mallory on Linkedin:
https://linkedin.com/mallorymejias
🤖 Please note this transcript was generated using (you guessed it) AI, so please excuse any errors 🤖
[00:00:00:14 - 00:00:09:17]
Mallory
Welcome to the Sidecar Sync Podcast, your home for all things innovation, artificial intelligence and associations.
[00:00:09:17 - 00:00:24:11]
Amith
and the world of associations. My name is Amith Nagarajan.
[00:00:24:11 - 00:00:26:12]
Mallory
And my name is Mallory Mejias.
[00:00:26:12 - 00:00:41:08]
Amith
And we are here to help you out with all sorts of cool fundamentals of AI, aren't we Mallory? The fundamentals sure are important in any discipline, any field of endeavor. And here in the world of AI and associations,
[00:00:42:11 - 00:00:43:09]
Amith
they're like bedrock.
[00:00:44:11 - 00:01:06:02]
Mallory
Exactly. Fundamentals are important across the board, not just with AI, but with anything. And believe it or not, it's been back since early 2024 that we did a two-part series on the fundamentals of AI. And still to this day, people tell us they listen back to that episode or they send friends that episode knowing, well, at two years old at this point, it's probably a bit outdated, some of it, not all of it.
[00:01:06:02 - 00:01:13:13]
Amith
I mean, not much has changed in that time, right? It's AI, it's kind of just poking along real slow, just kind of a chill progress.
[00:01:13:13 - 00:01:23:12]
Mallory
For sure. I feel like a lot has changed in terms of advancements, but the fundamentals themselves, some of them are the same. I mean, we'll get into that.
[00:01:23:12 - 00:02:46:08]
Amith
I mean, I think it's one of those things where the fundamentals of any field can change over time. What we think we know about nutrition is very different than what we thought we knew about nutrition 50 years ago. And the same can be said for many other fields. And so I think we have to be open-minded first and foremost when we're exploring a new space. And artificial intelligence is a 70-year-old pursuit. It's actually far older than that in concept, but as a practical matter, people have been actively exploring how to advance thinking machines of some sort for a little bit over 70 years now. And for the last 10 or 15 years, really the modern kind of evolution of that with deep learning, going back actually 15, 16 years exactly. But a lot's changed in that 15-year timeframe, enough to recast and rethink everything of what you can do, what you can't do. So people a lot of times tell me, "Well, but with AI, no, I can't do X, but I can do Y." I'm like, "Well, but do you really want to do Y? And are you sure you can't do X?" So it's questioning those things. So fundamentals are like bedrock, unfortunately, in the world of AI because things are changing so fast, you have to reconsider what actually, what's it built on. And so I'm super excited that we're doing this again, Mallory, because it is one of the fundamental elements of the Sidecar Sync podcast, those evergreen episodes, this one included.
[00:02:46:08 - 00:03:02:23]
Mallory
Yep. For the listener that has been with us since the very beginning or one of the early episodes with the meets, who feels like they've got a pretty decent grasp maybe of AI as a concept, especially as it pertains to associations, do you think they should still listen to this fundamentals series that we're doing?
[00:03:02:23 - 00:03:28:09]
Amith
Totally. No matter how long you've been practicing something, I think it's great sometimes to get a fresh look at the basics. I'll give you an example from totally outside of this realm. I love skiing, and I've been skiing since I was a little kid, and I'm reasonably good at it. It's the only sport that I can barely stay balanced in, mainly because I've been just doing it for so long because I can overcome my natural inathletic ability to do this.
[00:03:29:11 - 00:04:41:05]
Amith
But every so often, first of all, I love skiing with really good skiers and seeing how they ski and what they do. But every once in a while, groups that I'm with will hire a guide for the day to take a surround resort mainly to get around lines. We'll get a group of four, six, or eight people, and we'll hire a ski instructor so we can skip the lines for the day. If you have a good-sized group of people, it isn't so crazy expensive to do that. But anyway, every once in a while, one of these instructors actually says, "Hey, I'm hired to be a ski instructor. I want to teach you something." And so they look at me and they go, "Look at you. You kind of think like you know how to ski, but you're doing all this wrong." And so I'm like, "All right, let me get my ego out over here and listen to what this person is saying." And you do learn fundamental things, and you're like, "Oh, wait a second. Yeah, that actually works." Or, for example, in the world of skiing, the way you might have skied in the 1980s or in the 1990s, it's changed a lot because the equipment has changed a lot. So gravity hasn't changed. I haven't really changed all that much, but the equipment has changed a whole bunch. So you have to stop and think and kind of reflect a little bit. So with AI, instead of every maybe 30 years, like with skiing, you might want to think about it maybe every 30 months.
[00:04:41:05 - 00:06:24:11]
Mallory
Yep, maybe every six months. So you said it here, Amith. If you're listening to this and you feel like, "I've got a pretty decent grasp on artificial intelligence," you don't want to be like a meath on the skis with the instructor thinking, "I know what I'm doing," and having them point out different things. So stay tuned. We've got an exciting first episode. This one's going to be very fundamental. What is AI, generative AI? We're going to get to what's happening under the hood, not getting too technical, but just so you have an idea of when we talk generative AI, what we're speaking about. So starting off with the fact that, Amith, you kind of teased up my point already, AI is not new. Generative AI is the thing that has kind of exploded in recent times. So let's go back to where most of us started. When chatGBT opened to the public in late 2022, you didn't need to be a developer or buy anything or understand a single thing about how it worked. You typed into a box in your browser and something wrote back. It had 100 million users in about two months, faster than any consumer product before it. And the reasonable conclusion for almost everybody was that AI had just arrived. But here's the thing a lot of people may not know. AI is not new. As a field, it goes back about seven years, like you said, roughly to the dawn of the digital computer. And many of the ideas we're using today were worked out decades ago. It just didn't work very well for most of that stretch. And honestly, you'd already been using AI probably for years without calling it that. Every time your spam filter caught something or a streaming service picked your next show, that was narrow AI. Very good at exactly one job and useless outside of that. So my question for you, Amith, is if this field is 70 years old, why did it feel like AI suddenly popped into the world a few years ago?
[00:06:24:11 - 00:08:31:01]
Amith
Well, I think it's the convergence of a number of things. So first and foremost is compute. We had very limited computing capability, digital computing capability 70 years ago. We had lots of ideas. We were kind of at the beginning of that digital age. It was very exciting. But the compute we had collectively as a world back in 1950 was probably equal to less than the smartphone in your pocket, probably dramatically less than that, right? It'd be interesting to kind of sum that up and see what it would be, but it would probably be a lot less than one thousandth of what your phone can do. So we had very limited computing resources. With limited compute, it is hard to do even deterministic or rules based computation where you're doing not AI, but just having a computer literally follow instructions step one, step two, step three. But to do what we're talking about, it's very compute intensive. The other thing is that as we've struggled with the lack of compute, we've also struggled with the lack of data availability. Up until fairly recently, we haven't had a central resource where it's possible to get amounts of data in one place. Well, enter the World Wide Web and the scaling of the Web and the availability of a ton of content of a wide variety of types, including, of course, text, but also code, also video and audio, all these different pieces of content that are out there. That was a second component that was required. And then algorithmically, many of the ideas that are now actually working were thought of decades and decades ago by a lot of different people and actually were somewhat invented independently at different times by different people. But because of the absence of sufficient compute or sufficient data, they always seem to be wrong. People thought, "Oh, neural networking, that's just a cool sci-fi idea, but it's never going to work." And this is as recently as the 90s and 2000s, right? And so it took until the early 2010s for advances in all of those areas to be sufficient
[00:08:33:11 - 00:08:52:15]
Amith
to result in what we started to see as the deep learning revolution. So to me, that's what a lot of what it is. It's convergence more so than it is a brand new idea. People have been talking about intelligent machines for a very long time. And it's obviously so exciting that all of us here today get to live through this transition. It's also a little bit scary or a lot scary, but
[00:08:53:17 - 00:09:00:23]
Amith
it's here. And it didn't happen overnight. It was the combination and compounding of a whole lot of fields of progress coming together.
[00:09:02:03 - 00:09:08:07]
Mallory
So it sounds like the compute and the data were necessary to bring those ideas to life over the past 70 years.
[00:09:08:07 - 00:09:56:15]
Amith
For sure. And to be clear, it's not that computer scientists just have been resting on their laurels algorithmically since 70 years ago or anything like that. There's been tons of research and tons of advancements with all sorts of different ideas for the computation, like what do you do with all those resources? And back in 2017, there was a remarkable paper that's now known as the Transformers paper that was published out of Google that really did lay the foundation for modern large language models. And it took a number of years for that to really get attention. But it was ultimately the piece that resulted in the general purpose part of what you described earlier. We've had smart AI models in narrow spaces for a long time. The Transformers were the first architectural innovation that really made it possible to train a wider, more general AI.
[00:09:58:06 - 00:10:29:10]
Mallory
So if back in 2022 you were testing out chat GPT, you were using generative AI. So it's the kind that doesn't just sort or recognize but makes new things, text, images, audio, video. That's the branch that exploded, not necessarily AI as a whole. It's also worth naming because this is something you may have heard out in the ether, the term artificial general intelligence. This is kind of the version of AI you might see in the movies that can do anything a person can. That one doesn't exist dot, dot, dot yet.
[00:10:30:18 - 00:10:50:06]
Mallory
We'll talk about that in just a second. But I mean, I feel like some people say AI is overhyped. And then you've got the other school of people that say it's the most transformative tech of our lifetimes. And maybe both of those things can be true at the same time. But I'm curious how I think if you listen to the podcast, you probably know where Amith falls. But where do you fall along those two things? Or do you agree with both?
[00:10:50:06 - 00:12:48:19]
Amith
Well, you can agree with both at the same time, I think. I do think it's massively overhyped, but it's a timing issue. And so, you know, some things can be massively overhyped in the short term relative to the value they're creating in the world versus being massively underhyped or underappreciated relative to their long term potential. Electricity is a good example of that, right? Telecommunications are another good example of that. Digital computing is another good example of that. These things were all transformative, world changing, general purpose technologies. When they initially came on the scene, there was incredible enthusiasm. But then very quickly, people were saying, oh, wait a second. I'm not seeing, you know, the actual results. Where is it on my P&L that I actually got benefit from technology X, plug-in, plug-in, whatever you want? And that's exactly what's happening with AI right now. Corporations and nonprofits alike are saying, you know, where's the meat? You know, where am I going to get the actual results from when it comes to AI? And it's a damn good question to be asking because there's so many projects and so many attempts and frankly, not that many results to show for it yet. I do think there are results, obviously, out there. We're very excited about many of the things we talk about here on the Sidecar Sync podcast. But broadly speaking, relative to the level of investment the world has put into AI, even in the last two or three years, the level of GDP growth corresponding with that investment is still de minimis. So that is an important thing to be considered of. At the same time, the potential for AI is greater than all other transformative technologies of the past combined and then some. So I think it's under hyped actually relative to its transformative potential, largely because it's hard for us, myself included, all of us, to really envision what the world looks like with unlimited, abundant and largely free intelligence. And by the way, is AGI really not here yet? What do you think, Mallory?
[00:12:48:19 - 00:13:01:02]
Mallory
I mean, I feel like technically it's not here yet, right? Wouldn't we need fully functional world models, which maybe you can argue already exist for AGI to be here? What do you think?
[00:13:03:13 - 00:13:23:15]
Amith
I don't know. You know, it's one of these things where like the original definitions people were throwing around is a computer that's capable of doing all general purpose tasks at kind of a midpoint of capability, right? So it's like the median, right? If you think about a bell curve of capability, so it's not better than every lawyer and every accountant and every
[00:13:25:15 - 00:13:58:05]
Amith
nurse on earth, but better than the average one. That was originally what a lot of people were saying AGI is or competent enough to actually be in the field was even an earlier definition, which isn't even necessarily at the average point. I would argue that for white collar labor, things that don't require the physical world, AI is, I think, pretty much there. I think it's kind of hard to argue with the data and the experiences we're having. Now, it's not capable of being plugged into everything yet, but is the raw brainpower of Opus 5 or Fable or GPT 5.6 or KIMI K3
[00:13:59:08 - 00:14:16:17]
Amith
on par with the average human and the average white collar job? I think the answer is a definite yes. And so that doesn't make me say, "Okay, like, you know, AGI is here, game's over for all of us." I don't think that's the case at all. In fact, I've never believed that to be true.
[00:14:17:19 - 00:14:23:09]
Amith
But my point is that I think it's a moving goalpost is really what I'm trying to say. Now people are saying it's
[00:14:24:19 - 00:14:55:00]
Amith
more competent than any human at all economically valuable tasks, right? It's a much broader definition and a much higher bar. I think one of the reasons people are moving that bar is because the value hasn't yet been accrued to the economy and by extension your association or businesses, at least not to the extent that the investment has suggested it should be. There's a lot to unpack in that statement, but I guess the point I would make is if you say you're there, then people are going to say, "Where's my results?"
[00:14:55:00 - 00:14:57:08]
Mallory
Right. You make a good point.
[00:14:58:08 - 00:14:59:11]
Amith
That's one of the reasons people are moving it.
[00:14:59:11 - 00:15:16:05]
Mallory
I was thinking more navigating the physical world in the way that I interact with a human. "Hey, could you go to the grocery store for me and drive this car there?" Which I guess AI could do. And then get the groceries and then come back and then cook them and anything that a human could do. But maybe that is a moving goalpost from where we were even just in 2022.
[00:15:16:05 - 00:15:59:08]
Amith
Once you get into the physical world, we clearly have a ways to go. And that does have to do with what you briefly mentioned with world models and vision and all that stuff. And that's not that far behind, but it's behind where we are with some other modalities. I just think it's an interesting thing to ask ourselves because we just kind of accept what the Frontier Labs are saying about AGI being a ways out there. I think the idea of artificial superintelligence is kind of the next term beyond AGI that a lot of people are focusing on. It's potentially interesting, right? That's going beyond all humans combined is typically how that's defined. You take all of our intellects combined, all eight billion of us, and you make one AI system that's smarter than all of us collectively. We're definitely not there yet. I don't know that we want to be there, but that's...
[00:15:59:08 - 00:16:00:09]
Mallory
I don't think we want to be there, Ami.
[00:16:00:09 - 00:16:05:19]
Amith
Yeah, that seems to be where things are heading, at least in the minds of some of the leaders in the Frontier Labs.
[00:16:05:19 - 00:17:05:20]
Mallory
So we're going to have artificial super general intelligence, artificial extra super general intelligence. You'll learn in these episodes. There's some fun terminology. But I want to move to a little bit of what's happening under the hood. The engine behind almost all of this that we're talking about is the neural network. And as Ami talked about, this is software loosely modeled on the way neurons in your past signals to one another. The surprising part is that nobody handwrites the rules. The model learns them itself during a phase called training, which can be a long and expensive process of finding patterns across an enormous amount of data. Using the finished model is a completely separate phase called inference. So the moment you type something in and it answers, you can think of training as the schooling and then inference is actually working the job. It's worth keeping straight because training happens once and costs a lot of money while inference happens every time you hit enter. Ami, we're getting into some terminology. What do you think the association leader listening needs to know about inference?
[00:17:05:20 - 00:17:40:24]
Amith
Yeah, I mean, I think you really characterized it. Well, Mallory, it's it's it's what you do when you use chat GPT or Claude or any other AI tool. Your use of the model is inference. That's all it is. It's basically a process where the AI model that is all the way already been trained takes a request in and produces an answer. That's called inference, whereas the process of creating the model initially is called training, just like you might take a young person and put them through their schooling and you ultimately have a college grad or a high school grad or a PhD.
[00:17:42:00 - 00:17:51:06]
Amith
That process is the training and then they enter the workforce and they start taking on task after task after task. Each task in that analogy would be the inference process.
[00:17:51:06 - 00:17:55:09]
Mallory
And you can choose as a leader your inference platform, right?
[00:17:55:09 - 00:21:38:24]
Amith
Sure can. Yeah. And there's a lot of choice. This is another thing that's a big update from where we were even two years ago when we were talking about this stuff in mid 2024. At that point, we were roughly what was it? A year and a half post chat GPT moment. A little bit longer than that, but not much. I think we were still in the GPT 4 era, maybe 4.1. And we were all this is pre reasoning models. This is where models were just basically picking their best first guess at every question. Now these models iterate and refine their answers, which has made them dramatically better, by the way. But I would say to you in that couple of year time frame, we had fewer choices because there was open AI. Claude really wasn't that popular two years ago. Claude existed. They'd been around, I think, since 2021 or 2022, but they hadn't really gotten their big hits yet with Claude code or even the Claude desktop app. They were starting to attract some users, but their stuff was generally quite behind open AI. Open AI was pretty much the undisputed leader in AI. So if you wanted the best AI, you just signed up for an open AI API account and built your stuff directly around them and good for them. But that advantage didn't last. There's tons of other choices that are just as good or in some cases considerably better. And not only are there many inference providers that are themselves model developers like open AI and thropic, which makes Claude, Google, which makes Gemini, these people are model developers. So they build the model, but they also provide inference. Think of it this way. You're getting the car from the auto manufacturer. That's like the model developer, but you're filling it up at a gas station or charging it. You know, if you have an electric car somewhere where you're getting your electricity from. And so the inference is kind of like that. You know, what you're doing day to day, running the car, driving the car, you need gas for it. So the inference providers are providing that fuel, so to speak. And Google and entropic and open AI are both. They're both model developers and inference providers. And then there are many other companies out there who are just inference providers and they will take they will either partner sometimes with some of the closed model labs like the ones I mentioned. But commonly what they'll do is they'll take open source, open weights models, which are free to use anywhere in the world by anyone and they will host them on their hardware. So a good example of that is a company called Fireworks here in the U.S. They're one of the larger ones. Cerebrus is a fast inference provider that has their own hardware platform. They provide inference services on top of their platform at very high rates of speed. They use things like Google's Gemma. They utilize a number of other models from from other providers. So you have so much choice. So to your earlier point, you can pick from a wide array of providers. This matters to you as an association for lots of reasons. But the one I want to hammer home is flexibility. This world is moving so crazy fast that you cannot predict. I don't care who you are or what seat you're in or how much you know about A.I. or how little you know about A.I. You cannot predict with any reasonable degree of accuracy who the leader is going to be in A.I. at any moment in time, even three months from now, much less three years from now. And last time I checked, associations like to build systems that last a little bit longer than a few months. So what you need to do is have flexibility in the way you run your A.I. workloads. So that's a key point to make. And there are so many inference providers. It makes sense to be able to take advantage of all of them. So that's another thing we can unpack later. But it's really, really important to know this, just fundamentally, that you have a lot of choice. You're not limited to just one or two things.
[00:21:38:24 - 00:22:26:00]
Mallory
Yep. Double clicking on inferences where you're running your models and you've got optionality and you should build in flexibility and how you're working with A.I. Now, the specific engine behind the chat interfaces everyone is using, like chat, GPT or Claude, like Amith mentioned, or Gemini is the large language model. It's something you might have heard referred to as an LLM. At its core, it's doing something pretty simple. Well, maybe not that simple. But when you talk about it in this way, it seems that way. Predicting the next chunk of language one piece at a time based on everything that came before. That's the whole trick. And it turns out to be enough to write, summarize, translate and hold a conversation that feels real. Amith, when people say A.I. is fancy autocomplete, are they wrong or is that a fair comparison to make?
[00:22:26:00 - 00:24:29:21]
Amith
It might have been technically somewhat reasonable to say that back in 2020, 21 and even 2022. But these model architectures have gone so far beyond that basic concept that it's really ridiculous to say that now. Even back then to say that is kind of ridiculous because, you know, it's like saying, OK, well, you know, I have a paper airplane in my hand. And so I understand the concept of flight. And I want to compare that to an Airbus A380, you know, Jumbo Jet. And so, you know, yes, they both take advantage of the same principles in terms of physics, but they're very different concepts. So, you know, oh, the A380 is just like a big paper airplane. That's all it is. Kind of ridiculous. The utility from LLMs, even back before a chat GPT with GPT-2 were tremendous compared to simplistic autocomplete systems. So I always found that to be a bit of a cop out. A lot of times you hear that when a new technology is making its way around, because people are like, I just want to dismiss that. I don't want to deal with it. That's a lot of times why you hear that. So I think it was definitely the wrong way to frame it, even if it was technically close to accurate originally, it's still the wrong way to frame it because the value in it and the applications of it are so dramatically different than autocomplete. But nowadays, it's not even a conversation because now these models not only predict the next chunk of text or predict the next pixel, the next word, the next piece of code, but they then reason over it. So they then stop themselves and say, hmm, is that right? Did I get that right? Oh, let me back up. Let me fix that. It's actually wrong. It's this. Oh, wait a second. Should I verify that information? I'm not sure if that's correct. Let me go and use the Google search or use some other proprietary source of content in the case of associations and make certain that that content that I'm about to generate and provide is right. So these models are far, far more sophisticated than that basic initial model. Of course, that's how everything works, right? You start off like web browsers back in 1993 were pretty simplistic compared to what they are now.
[00:24:29:21 - 00:24:44:22]
Mallory
I think the paper airplane that makes sense. It does. We call them language models, but these tools can look at a photo, listen to audio, watch video, and then also create those things too. What does multimodal mean in the context of AI models?
[00:24:44:22 - 00:25:46:23]
Amith
Yeah, the language and large language model is a bit of a misnomer at this point, as you just pointed out, Mallory. And I think it's best to just call them large models or just models, you know, because these models, the large part came from the size and we'll talk more about this at some point in terms of parameter size and weights and how big these things are in terms of the math that they're doing. But what matters is really that these are models that can process inputs in a variety of different formats. So you can pass in text. You can also pass in images. You can also send in audio and video or code. And, you know, the modalities that the models can natively take in is an important thing to be aware of because if the model can directly see an image, then it has more richness than simply a description of the image. You know, so if you were to take a one sentence summary of this podcast and share it with someone, I would certainly hope that listening to the entire podcast Mallory would be more valuable than just the one sentence.
[00:25:46:23 - 00:25:47:18]
Mallory
It better be.
[00:25:47:18 - 00:25:49:00]
Amith
Better be, right?
[00:25:50:01 - 00:28:09:18]
Amith
Yeah. And so I think that that's a good example of like, hey, if I take the caption off of an image and reason over that, I'm not going to be able to do as much than if I see the image. Similarly, if I listen to someone speak, I'm going to get a lot more out of them than if I am just reading the words that they spoke, right? So that's the difference between transcription and like reducing all these modalities to text, which is what early models tried to do versus now. The way they're working is, is that most of these frontier models are multimodal in terms of their inputs and some of them are also multimodal in terms of their outputs. So they can not only generate texts, but they can generate images. Some of them can generate audio and video natively. And so that's the key is that it's the difference between having a translator that says, oh, I see a picture of a golden retriever on the floor by a meets feet, which by the way, I have one right here on the floor and he's six months old as of today. And we love him. He's very chill, but that is a picture I painted in your brains with my words. But if I actually showed you a picture of my dog by my feet, that's obviously a lot richer and more descriptive. Same thing is true for AI. If you actually feed the image to the AI, it can do a lot more. So earlier models couldn't really do this. Now you kind of take it for granted. I would actually argue that many of the users I talk to, whether they're CEOs or membership directors or technologists, oftentimes are like, oh, yeah, yeah, yeah. We kind of know that it can, you can include anything, but we're just, we're just working with text right now. Like, well, but I get it, but that's like, I know television is out there. I'm going to stick with radio. Um, or I know television and radio are out there, but I'm going to stick with the newspaper. It's like each modality has its pros and cons, but not experimenting with them means you can't really, you can't really tell someone, Hey, this is what it's like to be on a roller coaster. Until you actually have been on a roller coaster. It's really hard to kind of relate to that experience. Similarly, as a user, if you haven't felt frankly, kind of the magic of being able to put images in and get images out, same thing with video and audio or trying like audio models, like real time audio, you haven't really experienced it. And you haven't opened up your own lens wide enough to understand what the possibilities are. So multimodality, I think is a big, big deal. Most people who are even say they're heavy AI users are primarily text based.
[00:28:09:18 - 00:28:31:01]
Mallory
And it's funny because like you said, I mean, people would argue, well, text is just how we do everything. It's how we communicate. It's how we work in business. But if you think about humanity, it's, it's audio, it's talking to each other. It's visual. It's looking at body language. So I think we've programmed ourselves to work with so much text because that's what worked at the time. But I think moving forward, it'll be interesting to see those modalities expand.
[00:28:31:01 - 00:30:13:03]
Amith
Totally. Yeah. And it's, it's so interesting too, because, um, you know, we started out not very long ago, just being able to try to do things actually, not with text modality, the big breakthrough in AI that led to deep learning was with images. It was around this thing called ImageNet, a competition we have talked about here on the pod before, but the quick recap of that is, um, back in the 2000s, uh, there was an effort to collect, uh, enough images to train AI models on so that they could auto classify, auto detect what's in the image and provide, you know, reasonable descriptions of them. And this challenge was called ImageNet and for a long, long time models, you know, were dismal failures at solving this problem until, um, early 2010s, we had an introduction of a neural net architecture that was actually trained on earlier Nvidia GPUs, uh, called AlexNet, which is introduced by the university of Toronto. And this one by leaps and bounds that particular year and he used a novel architecture that was based on neural networks. And that led to this whole deep learning revolution and that was in the image space. And so, um, and, and back then it was considered magical that you could show a computer, an image of a dog and have it correctly say, this is a dog, much less, this is a golden retriever. That seems like it might weigh 50 pounds or 80 pounds and it might be a certain age. None of that was really a thing. It was just dog, cat, person, table, right? And even that our computers were horrible at it. So that's where we've come in, in 16 years. Um, so the point is, is that images are really a rich modality and that's just one of the additional modalities that I really would urge folks to start exploring using with, with their AIs.
[00:30:14:12 - 00:30:33:11]
Mallory
I mean, you briefly touched on the idea, I think of open weights models when you were talking about inference, but I fear it may have gotten lost in that. And we do spend a lot of time on this podcast talking about open models and closed models. Could you give us a high level for a listener that's hearing that for the first time? How would you describe that? An open weights AI model versus a closed one?
[00:30:33:11 - 00:33:03:15]
Amith
Yeah. So open versus closed simply means kind of what it sounds like. The closed models are the ones that, um, the only, only the creator has the keys or has the weights and only the creator has the ability to run the model. So remember there's the training process and there's the inference process, the training process. Some typically large company with a lot of resources is going to train these models and they can choose to either keep it to themselves. And that's what the closed model is, or they can choose to share it with the world and say, here you go. And that's what open means. And so the open weights models are ones where people will publish the model and anybody can download the model with its weights. The weights are really the heart of what the model is. And then you can run it on any computer, but you have to have a lot of, a lot of compute for it typically, but open weights model. Are all about, it's very similar to the ethos of open source software, the same types of, the same idea behind Linux, behind WordPress, behind many of the software applications that you use day to day, even without knowing it are open source. In the world of AI, the source code actually isn't all that important and it's usually fairly small. It's the weights of the model, which is really kind of, uh, the gray matter. It's the substance of what's in the AI brain is what the weights are. And so when you, when you share those, that means anyone else can run it. Um, so open AI is in spite of the name, not an open, uh, weights company. They started off that way, but now they are very much a closed model company. Anthropic has always been a closed model company. Google is both. Google has the Gemini model series, which they do not release weights for, but they also have the Gemma series, which is G-E-M-M-A. Uh, the Gemma series, um, is their open weights offering. And it's a little bit less powerful than Gemini, but I actually kind of have to hand it to Google because like Gemma 4 is far more powerful than the prior version of Gemini. They're closed, uh, proprietary models. So they're two different business models. The argument in favor of closed for people who are building that way is you get to control the inference. So you, it's like having a toll road. You're keeping the toll road to yourself and you're charging a fee anytime someone wants to drive on it. Whereas open weights is you're giving away the technology to build the road. And anybody can put a road anywhere they want and you don't get to monetize the traffic on it. But the idea is, is that those companies see a pathway to making money, doing other things, either providing services or applications on top of it, or there might be other motivations outside of profit.
[00:33:05:04 - 00:33:09:12]
Mallory
You said weights are the heart of the model, but can you tell us what weights actually are? Sure.
[00:33:10:13 - 00:35:28:06]
Amith
So if we zoom out and we have to take some liberties here and compare artificial neural networks with biological neural networks and the basic idea in the biological realm, which I'll say conveys essentially to the artificial realm is the idea of a neuron as being kind of this basic unit of compute. It's basically this, this thing that can fire, right? And neurons firing, they fire in a way where essentially think of them as being connected to other neurons so that the neurons, when they fire, they fire across these things called synapses. And the synapses are essentially like the wires that connect the neurons and the neurons can cluster together to form functionality. So in a biological neural network, aka a brain, you have these clusters that do things, like, for example, clusters of neurons that are wired together that are responsible for sight or smell or for hearing. And obviously, these are incredible over simplifications of things. Memory systems in the world of biological neural networks or brains are also intertwined in the same concept. They're part of that same neural network. In the world of artificial neural networks, you also have the concept of neurons and synapses, but the synapses are weights. So we call weights basically means it's the strength of the connections between neurons. So in a biological neural network, if I have a very strong synapse, which is like physically strong between two neurons, that means they're firing together a lot. And so necessarily then to support that additional electrical current, the synapse is stronger, heavier, thicker, right? And it's just more robust between those neurons that they fire together. A friend of mine likes to say that the neurons that fire together wire together. So it's a clean way to remember it. And so similarly in the world of artificial neural networks, it's the same idea. The neurons that are most closely connected to each other have stronger weights. The way that you have X number of billions or even trillions of parameters, it's the actual values in this giant file called the weights that basically define what these weights are between the different neurons. That's essentially what these neural networks are, is a massive file with a whole bunch of numbers in them.
[00:35:28:06 - 00:35:30:17]
Mallory
Sounds like a lot of fun.
[00:35:30:17 - 00:35:46:02]
Amith
Yeah, it sounds like fun to some. Not so much to me, but it's interesting. But it's kind of like about as decipherable as if you looked at a brain. It looks like a jumbled up bunch of mush, right? And that's kind of what these files look like if you were to open them.
[00:35:47:12 - 00:36:03:13]
Amith
But when you load them into this frankly very minimal amount of software, which is the model, the code of the model, and then you fire it up, you have a working neural network. And it's actually kind of amazing that it works. So it's pretty powerful.
[00:36:03:13 - 00:36:35:14]
Mallory
we've covered quite a bit in part one of this Foundations of AI series. One, AI is about 70 years old and it's generative AI that exploded in 2022. Neural networks learn their own rules rather than being programmed. And the weights are the substance of what a model actually knows. And we also talked about how you've got real choice in models and inference providers so it's important for you to build for flexibility. In the next part of the series, we're talking about how the cost of intelligence is getting cut in half
[00:36:35:14 - 00:36:41:03]
(Music Playing)
[00:36:51:21 - 00:37:08:20]
Mallory
Thanks for tuning into the Sidecar Sync podcast. If you want to dive deeper into anything mentioned in this episode, please check out the links in our show notes. And if you're looking for more in-depth AI education for you, your entire team, or your members, head to sidecar.ai.
[00:37:08:20 - 00:37:12:01]
(Music Playing)