Skip to content

Podcast

Introducing Maturity Maps — A New Way to Measure AI Adoption

The AI Daily Brief: Artificial Intelligence News and Analysis

Source ↗ ← All highlights
  • Why Internal AI Wins Can Still Mean Falling Behind
    • AI adoption without benchmarks can create false confidence because internal gains may still trail competitors.
    • Nathaniel Whittemore uses a marketing example where 30% output growth sounds strong until peers are actually growing 50%. Transcript: Nathaniel Whittemore Now, one of the things that I’ve been thinking about a lot over the last six months or so is just how much we need a totally different set of data and benchmarks for this new AI era. Everyone is adapting incredibly quickly right now, or at least they’re trying to. It’s new processes, new workflows, new tooling, new everything. And by and large, we’re doing all that exploration without a map. Let me give you a practical example of where I think our lack of benchmarks could actually very significantly and meaningfully negatively impact a company when it comes to their AI Adoption. Let’s say that you are an early adopter company. Across your different functions, you’ve had really strong hands-on efforts to get your AI up and running. In the absence of knowing exactly what to measure, you’re just trying to measure whatever you can, and early results are pretty positive. For example, in the marketing function, you have increased your content output 30% year over year without any sort of proportional increase in the resources it takes to produce that Content. Now that 30% year over year growth sounds great, but what if I told you that all of your competitors had actually grown their content output by 50% In AI world, this is actually not a far-fetched Scenario, and it shows how the need for better benchmarks and numbers is not just a vanity exercise. When we don’t know how we’re doing relative to peers and competitors, it makes it really hard for us to judge what we need to change, what we need to shift, and what we need to do next. (Time 0:01:06)
  • Capability Overhang Makes Old AI Benchmarks Obsolete
    • Raw model capability is not the real bottleneck anymore; the limiting factor is the systems around AI that turn capability into value.
    • Nathaniel Whittemore argues traditional vendor-ranking frameworks like Magic Quadrants miss how AI adoption actually works inside companies. Transcript: Nathaniel Whittemore Anyone who’s felt the sting of the capability overhang, in other words, the gap between what AI can do and what we’re actually using it for, knows that raw capability isn’t really the Question. It’s the systems we put around it to get value from it. And unfortunately, the research and information apparatus just has not adapted to this new reality. Not to pick on Gardner specifically, but they’re the biggest in the space, and so sort of present an easy target. Tried and true benchmarks and information products like Gardner’s Magic Quadrant have literally never been less useful than they are right now. The idea that success in something like AI application development was going to be even a little bit dictated by choosing the right AI application development platform vendor is just So far outside of the reality of these tools as to be almost actively harmful if that’s where you’re putting your time and effort when it comes to trying to figure out how to adopt AI. Now, Gartner is more than the Magic Quadrant, and they are doing lots to try to catch up to the AI world, so it’s not to single them out. It’s more to make the point that we are in desperate need of some new frameworks, some new benchmarks, and some new tools. (Time 0:03:34)
  • The Six Dimensions That Define AI Maturity
    • Maturity Maps measure AI readiness across six dimensions instead of just counting use cases.
    • The framework scores deployment depth, systems integration, data, outcomes, people, and governance because change management and infrastructure determine whether use cases deliver value. Transcript: Nathaniel Whittemore When we’re doing AI readiness and planning assessments at Superintelligent, we’re not just thinking about what use cases a company should do, but what’s the full set of change management And infrastructure development and new policy and investment in people and all this other stuff needs to go around it to actually get value from those use cases. And that led to the development of the framework, which I’m going to be sharing today, which we call for simplicity, AI maturity maps. Now, the concept of maturity is certainly not some proprietary thing that we invented. Maturity is just a heuristic and a framework to look at where different organizations are around some key areas relative to one another and where they should be. So the way that maturity maps work is that they organize AI and agent maturity into six different categories. Those categories are first deployment depth, which is sort of an expanded notion of use cases. Deployment depth in the context of AI maturity not only thinks about how many use cases you have in play, but how much those use cases are assistance versus full workflow automations Versus actual applied agentic systems that are doing work with some meaningful degree of autonomy. The second category is systems integration. This is a measure of how deeply integrated the AI solutions and workflows you’re deploying are integrated with the existing systems that run your enterprise. Is everyone using ChatGPT independently, or does your CRM system have an agent running through it, automatically extracting insights, making recommendations, and even setting Up new outreach campaigns? Systems integration is in some ways one part of the measure of how good the context that an enterprise’s AI has to work with. Now, the other piece that relates to context is, of course, data. How much, what quality, and how well-managed is your company’s AI’s access to your company’s data? Does it require people dropping in PDFs? Do you have company knowledge all set up on MCP servers? How does the AI that your company is looking to transform your company have access to the information it needs to know what that transformation should look like? Outcomes is almost a measure of measurement. Are all of your deployments pilots at experiments, or do you have a track record of actual demonstrable and measured outcomes? Outcomes in some ways are the information you need to know what you should do next across all these other dimensions. The fifth dimension of AI maturity maps is people, and this is an admittedly broad category. A big part of this refers to upskilling and capabilities, but another piece has to do with attitudes. Given that one of the major barriers to adoption in many companies is not just going to be skills using AI, but attitudes towards AI, people is an extremely important and unfortunately, As we’ll see, often neglected piece of the AI maturity pie. Lastly, of course, is governance. How clear, how established, how communicable, how known are the rules and guidelines and access provisioning around your AI systems? Do people know where to go to get the permissions they need? Do they know what expectations are? When issues come up, are there mechanisms for resolving those issues? So those are the six areas across which we look at AI maturity. (Time 0:06:09)
  • On Track Means Where Companies Should Be
    • The maps separate average performance from the on-track line, which marks where firms should be rather than where they are.
    • Nathaniel Whittemore says most organizations sit below that line, making the chart a visualization of the capability overhang. Transcript: Nathaniel Whittemore What came out of that is the chart that you see here, which plots each of these six categories within a specific function on a five-point scale. Number three, the center of the chart, is the on-track line. In other words, where an average organization should be. And the word should, as you’ll see, is doing a lot of heavy lifting there. Now, if on-track is a three, four is ahead and five is significantly ahead, while two is behind and one is significantly behind, is the idea is that when you look at a maturity map, without Having to read a lot of words, you can instantly see the gaps between where organizations should be and where the average organization actually is, and when you compare your organization To it, also see where you are relative to both the general on-track line and the average. So clarifying this a little bit more, a quarterly’s designation of on-track is not where the average organization is. It is a subjective measure of where we think the average organization should be. As you’ll see when you dig into this quarter’s numbers, in the vast majority of cases, we believe that the average organization is behind that on-track line across pretty much all of These dimensions. To use a term that comes up a lot on this show, the fact that organizations tend to be behind this on-track line is effectively a visualization of the capability overhang. Now at this point you might be wondering, well what gives you authority to determine what the on-track line is? It’s a totally reasonable question and believe it or not, it is a little bit more at least than just my opinion. We have a few different places to pull from. The first is the sort of proprietary research and surveying that we do as part of AIDB Intel, which gives us some pretty good insight into where particularly leading organizations Are. Second, it’s super intelligent, given that we are doing thousands and thousands of voice agent interviews every month to help organizations assess their AI maturity and plan their AI strategy. That’s another pretty unique source of frontline data. And then combined with that, we built a system to go out and effectively aggregate pretty much every new survey or study that comes out that even vaguely touches AI. You might have heard me mention before that my most useful open clause are my research open clause, and this is one of the main things that they do. They are in a never-ending 24-hour constantly hunting loop to both surface new sources, to assess those sources in terms of their legitimacy, credibility, and bias, and then to integrate That information into our larger assessment system. (Time 0:09:38)
  • The Maturity Maps Rest On A Massive Research Base
    • Nathaniel Whittemore built the Q2 maps from a large evidence base rather than a few anecdotes or vendor claims.
    • The system aggregated 480-plus studies from one quarter, covering 150,000-plus respondents across more than 50 countries. Transcript: Nathaniel Whittemore And then combined with that, we built a system to go out and effectively aggregate pretty much every new survey or study that comes out that even vaguely touches AI. You might have heard me mention before that my most useful open clause are my research open clause, and this is one of the main things that they do. They are in a never-ending 24-hour constantly hunting loop to both surface new sources, to assess those sources in terms of their legitimacy, credibility, and bias, and then to integrate That information into our larger assessment system. There are more than 480 studies and surveys from the last quarter that went into these Q2 maturity maps. Among the sources that have explicit sample sizes, the combined survey respondent base exceeds 150,000 professionals across more than 50 countries. The types of source categories that we have are one, big foreign top-tier consulting firm research. There’s over 20 of those sources in that mix, major platform earnings and public market statements. Analyst firm predictions in research from companies like Gardner, Forrester, and IDC. Function-specific regular or annual surveys, such as Stack Overflow’s engineering study. Or other similar things for areas like marketing, legal, and IT. Academic and government research. Behavioral data sources, where companies that have access to some unique user behavior data aggregate, analyze, and share that. A good example of that is Jellyfish’s AI coding benchmark, which used behavioral data from more than 200,000 engineers across 700 companies with 20 million PRs. Finally, there are, of course, practitioner reports and vendor case studies, although the system is careful to rate them with some amount of skepticism given that they are, of course, Selling something. (Time 0:11:27)
  • AI Adoption Is Rising Faster Than Real Embedding
    • Enterprise AI shows an adoption embedding gap where companies claim usage but rarely deploy AI deeply into workflows.
    • People are the biggest neglected constraint: seven of ten functions scored significantly behind, and Deloitte found 93% of spend went to infrastructure versus 7% to people. Transcript: Nathaniel Whittemore So in Q2, what are some of the patterns that we saw? The first you might call the adoption embedding gap. Basically, every single function-specific survey reports the same pattern. High claimed adoption, but at fairly low depth and utilization. This is maybe the most dominant finding across all these sources, that the story of Q2 when it comes to enterprise AI adoption is not just the capability overhang in general, but the Applied capability overhang even when it comes to adoption inside an individual organization. A second very common finding across a huge array of these sources, there tends to be a fairly big gap between worker-level data and leader-level data. For example, one study found that in the area of customer service, 72% of leaders said that their AI training was adequate, with 55% of their employees disagreeing. In HR, a huge percentage of leaders report that AI is a priority, but more than two-thirds of HR staff say that their organizations are not proactive in upskilling. In fact, one can argue that people are the bottleneck that is not getting nearly enough investment. We gave seven of the 10 functions a score of one significantly behind when it came to that people category. The irony is that one could argue that the single largest barrier to converting AI adoption into AI value is on the human side, and it’s the thing organizations are spending the least On. To use one dramatic example, Deloitte found 93% of AI spend going to infrastructure, with only 7% going to anything related to people. (Time 0:15:36)
  • Data And ROI Measurement Are The Main Ceiling
    • Data access and outcome measurement now cap AI maturity across most functions.
    • Eight of ten functions scored 1 or 1.5 on data, and firms rushed into adoption before building ways to measure ROI. Transcript: Nathaniel Whittemore Now, outside of people, another finding across all of these sources is that data is kind of the ceiling on everything else. Eight of the ten functions score a 1 or a 1.5 on data. Now, I don’t need to beat this drum anymore for this audience, but obviously without proprietary context feeding AI, things like your code base, your customer history, your deal data, You really are not going to get past basic assisted usage no matter how good the assistant tools get. One could argue that data is not one pillar among six, but the floor constraint that caps all the others. Another area with universal challenge is around outcome measurement. And this is not surprising. Given how much pressure there has been to adopt AI as fast as possible, one of the consequences of that is that no one slowed down, paused their adoption while they went out and figured Out how to actually measure the ROI of all those investments. Can you imagine right now someone in the C-suite suggesting with a straight face that you take six months off of adoption to figure out better ways to measure the ROI first to ensure that You weren’t spending too much. That, my friends, is a recipe for an early retirement. The byproduct of that, however, is that the actual evidence for AI ROI is pretty thin. Now, I will say that if I had to make a prediction about one area where you’re going to see the biggest glow up this year, I think there are tons and tons of efforts around ROI measurement, And I would expect that to jump significantly in the quarters to come. (Time 0:17:05)
  • Customer Service Shows The Human Cost Of Poor AI Change
    • Customer service may preview what happens when companies automate fast without preparing workers for the new human role.
    • AI removes routine tickets, leaving agents harder emotional cases; Nathaniel Whittemore cites high stress, training gaps, and rising burnout. Transcript: Nathaniel Whittemore Now, a few observations looking across the different functions. While customer service did have a couple areas that we rated as on track, it also, I think, reveals something that could be a harbinger for other areas. Remember, when it came to CS, we heard 72% of leaders say that training is adequate, but 55% of people actually working in CS say it’s not. 87% of customer service workers report high stress, and 75% of leaders acknowledge that AI may be increasing stress. So you’ve got a situation where AI is absorbing routine cases, humans get the harder, more emotional ones, many people might not be trained for that shift, and that, plus the fact of Just increased questions about the long-term sustainability of your job, the result is stress, anxiety, and burnout. Basically, CS could be the canary in the coal mine for what happens when you deploy AI without investing simultaneously in the humans who work alongside it. One of the areas that I think is interesting to point out is the two rating or behind when it comes to governance. In most organizations, IT owns AI governance for a big part of, if not the entire organization. And yet only 54% have centralized frameworks. 50% of AI agents are unmonitored. 88% have had security incidents. And the question becomes, if the governance function is behind ungovernance, what does that tell us about the rest of the organization? One of the most interesting findings when it comes to sales is that it might be the cleanest example of the adoption mirage. 88% of sales teams say they use AI, but only 24% have it in their actual revenue workflows. A fair bit of adoption then, in other words. This is why we rate them behind in deployment depth. Much of the quote-unquote adoption is reps using ChatGPT in a separate browser tab for email drafts and call prep. Which is not a bad thing at all. It’s just not the level of automation and autonomy that I think sales organizations are hoping for. The autonomous SDR dream has not fully come to fruition yet, and I don’t think that most sales organizations have figured out the right integration and balance between humans and agents In the new sales working system. (Time 0:19:15)
  • Sales And Ops Reveal The Adoption Mirage
    • Sales and operations both overstate AI maturity because reported adoption often hides shallow usage.
    • Sales reps mostly use ChatGPT beside core tools, while operations often counts pre-GenAI automation from years ago as current AI investment. Transcript: Nathaniel Whittemore Of the adoption mirage. 88% of sales teams say they use AI, but only 24% have it in their actual revenue workflows. A fair bit of adoption then, in other words. This is why we rate them behind in deployment depth. Much of the quote-unquote adoption is reps using ChatGPT in a separate browser tab for email drafts and call prep. Which is not a bad thing at all. It’s just not the level of automation and autonomy that I think sales organizations are hoping for. The autonomous SDR dream has not fully come to fruition yet, and I don’t think that most sales organizations have figured out the right integration and balance between humans and agents In the new sales working system. The deployment depth score on operations, I think, is another really interesting one. In some ways, operations has had automatable functions longer than any other function in the enterprise. Think statistical forecasting, rules-based inventory management, predictive maintenance. That’s all stuff that predates this latest wave of Gen AI by a decade. What that means is that when 90% of operations teams say they’re investing in AI, it sounds impressive. But when you actually look at what that is, a lot of it is legacy optimization infrastructure that’s been running since, I don’t know, 2015. The Gen AI layer on top is often very thin, mostly asking AI questions about operational data and generating reports. In fact, one study found that only 23% of operations groups even have a formal AI strategy. Operations is the function where the distinction between old automation and new AI maturity is showing up as a real distinct challenge. (Time 0:20:33)
  • Finance May Benefit From Governing AI Before Scaling It
    • Finance may become a late-moving winner because it already knows how to govern risky systems even if deployment is weak today.
    • Nathaniel Whittemore notes finance is on track for governance due to compliance muscle memory, then asks if that lets it leapfrog later. Transcript: Nathaniel Whittemore Lastly, one more interesting stat, outside of the technical areas of engineering and IT and the long duration area of customer service, which has had so much emphasis on automation, Finance is the only other function to hit on-track on any pillar, and it does so on governance. Now, why do you think this might be? It’s because, of course, finance is operating in an area where governance is not optional. 69% of CFOs report advanced or established AI risk governance frameworks. Why? SOX compliance, audit trails, fiduciary duty, decades of regulatory muscle memory. Basically, finance already knew how to govern risky tools even before AI existed. Now don’t look at the rest of finance, where we rated it significantly behind in every other category. Basically, they know how to control AI haven’t figured out how to use it. What’s interesting will be whether over the next few quarters, we see this turn into a tortoise and a hare thing. In other words, when finance does figure out how to deploy, will they do it more safely and more effectively than functions that deployed first and governed later? And will that actually allow them to catapult at some point and jump out ahead of other functions who at the moment feel like they are farther when it comes to deployment depth. So this is the idea of maturity maps. They are, of course, very nascent. This is the first quarter where we’ve actually fully trotted them out. And in the spirit of getting feedback and seeing how useful these things can be, we’re putting up the ability on the super intelligent website, bsuper.ai, to actually go not only review All of these maturity maps, but do a short quiz that actually shows you where your organization is relative to both the on-track line and what we think is the average. And I do want to emphasize that this is not an assessment. This is not an audit. This is a quiz. It is an online quiz that’s 18 questions that is going to give you a very general idea of where you stand. We of course have ways to go deeper with your organization and get much better data to actually inform these things, but we want as many people as possible to have access to this to actually Help us figure out where these lines should be and how we should evolve the entire system. Obviously, I’ll have links to all of this in the show notes, but again, that’s going to be at bsuper.ai slash quiz. In terms of where we want to take this next, in addition to just continuously having better and more sources of data, probably the most glaring thing that stands out to me is that we’re Trying to argue for one on-track line and one average line across all different organization types. 10-person startup by the same on-track lines as a 10,000 enterprise. I obviously believe that there’s enough value that it’s worth putting it out, while acknowledging that where I’d like to go next with this is vastly more gradations in both the on-track And the average lines, organized by things like organization size. Industry is another obvious area where you might see some fairly significant differences, and while I think it is the right call to start with some very broad- high-level functions, Obviously most organizations are a lot more nuanced than just having these 10 clear departments. (Time 0:21:57)