There is a conversation that repeats itself constantly with association staff who use AI every day. They describe themselves as heavy users, and by any reasonable measure they are. They have paid accounts, they have favorite prompts, and they reach for a model several times a day without thinking about it.
Then you ask what they actually put into it, and the answer is always text, occasionally in the form of a document that was already text before they pasted it.
When the conversation turns to images, audio, video, or screen recordings, the response is usually some version of yes, we know that is possible, but we are working with text right now.
That answer makes complete sense. Association work runs on text and always has. Board packets, bylaws, member correspondence, grant applications, conference copy, and a great deal of email are the actual medium of the job, so text is the natural thing to reach for.
There is simply more available than most of us are using, and the distance between what these systems accept and what we tend to hand them may be the largest source of untapped value sitting inside associations right now. Closing it costs nothing.
Part of the problem is terminology. The systems behind the chat interfaces everyone uses are commonly called large language models, and that name has quietly stopped being accurate.
The word "large" came from size, referring to the sheer scale of the mathematics involved. The word "language" came from what these systems could originally handle, which was text and only text. That constraint is gone. Current frontier models accept text, images, audio, video, and code as native inputs, and a growing number produce images, audio, and video as outputs rather than only describing them.
Calling them language models made sense in 2021. At this point the more honest term is simply models, because the language part describes a limitation they no longer have.
Names shape expectations. When a tool is called a language model, bringing it language is the obvious move, and there is no particular reason anyone would think to revisit that assumption after the underlying capability quietly expanded.
Here is why this matters more than it sounds like it should.
When you describe an image to a model instead of showing it the image, you are handing over a summary and asking for analysis of the summary. The model reasons over your caption rather than over what is actually in the picture. Everything you did not think to mention is gone, including the things you would not have known to mention.
The same holds for audio. Reading a transcript of someone speaking gives you the words. Listening to the recording gives you the words plus pace, hesitation, emphasis, where the energy rose and where it flattened. A transcript of a member focus group tells you what was said. The recording tells you how it was said, and anyone who has sat through a focus group knows those are different pieces of information.
Reducing everything to text was what early models required. They could not see or hear, so the only path was to translate the world into words first and accept whatever was lost in translation. That constraint shaped a generation of working habits, and the habits outlasted the constraint by several years.
Think of the difference between reading a one-sentence summary of an hour-long recording and listening to the recording itself. Nobody would argue those two experiences are equivalent. A text-only workflow makes that same trade on every input, and because it happens silently, there is rarely a moment that prompts anyone to reconsider it.
It is worth remembering how the current era of AI actually began, because it did not begin with text.
In the 2000s, researchers assembled an enormous collection of labeled images so that models could be trained to identify what appeared in a photograph. The effort became a competition called ImageNet, and for years the results were poor. Models were genuinely bad at answering a question a small child answers instantly.
Then in the early 2010s, a neural network architecture out of the University of Toronto called AlexNet, trained on early Nvidia graphics processors, beat the field by an enormous margin. That result launched the deep learning revolution that everything since has been built on, and it happened in the image space rather than the language space.
It is also worth sitting with how low the bar was at the time. Being able to show a computer a photograph and have it correctly answer "dog" was considered close to magical. Nobody was asking for the breed, the approximate weight, or the likely age. The categories were dog, cat, person, and table, and even at that the systems were unreliable.
That was roughly fourteen years ago. The models available to your association today can look at a photograph and tell you the breed, estimate the animal's size and age, and describe what is happening in the background. The visual capability was never an afterthought. It was the starting point.
The strongest argument for text is that business genuinely runs on it, and for good reason. Contracts, policies, minutes, and member communications need to exist in a form that can be filed, searched, signed, and referred back to five years later. Text does that better than any other format, which is why professional life is built around it and why it deserves to stay there.
What is worth noticing is that this is a practical inheritance rather than a reflection of how people naturally communicate. Human exchange is largely spoken and visual. We read tone, faces, posture, and gesture, and we did all of that long before writing existed. Text became the backbone of professional work because it was the format that could be reliably stored and transmitted, and that constraint shaped a great deal of what now feels simply normal.
Consider a comparison. Preferring radio to television is entirely defensible, and so is preferring the newspaper to both. Each format has real advantages, and none is strictly better than the others. The catch is that a preference only becomes an informed one after you have spent time with the alternatives, and most of us have not yet spent that time with anything other than text.
That is the one real cost of staying entirely with text. Text is not the wrong tool for most of what associations do. Using it exclusively, though, turns a wide capability into a keyhole view, and a keyhole gives you no way to judge how much of the room you are missing.
Describing a roller coaster to someone who has never ridden one does not transfer the experience. Reading about what multimodal AI can do works the same way. Until you have put an image in and watched what came back, or spoken to a model in real time and heard it respond, the capability stays theoretical, and theoretical capabilities do not make it into anyone's workflow.
The barrier here is habit rather than budget or technical skill, which means the fix is a series of small substitutions.
None of these require a new vendor, a project plan, or board approval. They require noticing the moment when you are about to summarize something for a model and choosing to show it instead.
Building that fluency matters for a reason beyond your own efficiency. Once handing a model an image or a recording feels routine, a different category of question opens up, and it points outward toward your members.
Your members may not live in text either. A field inspector, a clinician moving between patients, a contractor on a job site, or a technician with both hands occupied is unlikely to stop and type a careful question into a form. Someone in that position could photograph what they are looking at and ask about it, or simply ask out loud and get an answer back. The session that addresses their exact problem may already exist in your conference video library, where the genuinely useful portion is probably ninety seconds long and currently undiscoverable.
The same logic reaches into credentialing and professional education, where a submitted work sample, a photograph of finished work, or a recorded demonstration can carry evidence of competence that a written description cannot. It reaches into accessibility, where offering content in more than one format widens who can actually use what you produce.
What you need is enough hands-on familiarity with these formats to recognize the opportunity when it appears, rather than reaching for the text-shaped version of every member interaction because text is the shape you have practiced.
That is the real argument for experimenting now, while the stakes are low and the only thing at risk is ten minutes of your afternoon. The associations that get the most out of AI over the next few years will not necessarily be the ones with the biggest technology budgets. In many cases they will be the ones whose staff built the instinct to hand these systems the actual thing rather than a description of it, and then noticed where that same instinct could serve the people they exist to support.