Six Questions to Ask Before Investing in AI
Check the business need, alternatives, data, responsibilities, results, and long-term cost.
Linda Apsley
Managing Partner & CEO
Articles and videos from Linda and Allan on what to build, what it costs, and how teams use and maintain it.
Allan Carroll with Zhen Lu, CEO and co-founder of Runpod
32 min
Allan and Zhen Lu compare what they see at two layers of the AI stack: the infrastructure that runs models and the controls a business needs before agents can do real work. They cover where data should live, when open and customized models matter, replaceable model choices, evaluation that catches regressions, and human approval when agents act quickly.
Chapters
Lightly edited for readability.
The video opens with a clip from 16:29.
Zhen Lu: If they depend on somebody else's model, it can be taken down at any time, or it can be deprecated, for whatever reason. That's not great if all of the systems you rely on break because of something you have no control over. I think that's a pretty decent reason for customers that need to run business-critical workloads to run them on open models, on infrastructure they control.
Allan Carroll: Good to be here with you today, Zhen, to chat about what's going on in AI, what you're seeing at Runpod, and in your explorations out in the AI world. I'd love to hear a little about yourself, where you're coming from, and what you're thinking right now.
Zhen Lu: Thanks so much for having me, Allan. At Runpod, we're building the foundational platform for developers to build and run AI systems at scale. It's incredibly exciting. I see the foundations of software development changing, so it's an exciting time to be alive, and I'm looking forward to the conversation.
Allan Carroll: I know you're deep in the weeds with a lot of really interesting work. You're providing the foundational compute and a lot of the tools for it. We're seeing that too. I've been working in AI, building, learning, and studying, for about 20 years, and now building AI products at StrataEdge. We're lucky enough to work with a lot of big companies that are putting AI into their internal tools and workflows and learning how to use it themselves. We see what it looks like on the ground for people who use it day to day and build internal and external tools around it.
Zhen Lu: It's a great complement. I see StrataEdge starting at the enterprise, and Runpod started with developers and community. In the next year or two, we'll probably meet in the middle somewhere, because everybody needs to be an expert in this going forward.
Allan Carroll: I think you're right. One thing I've been thinking about, watching Runpod, is the insatiable appetite for compute and intelligence. I just don't see it flattening. How do you think companies will eventually manage the costs and their use of compute? Or are you seeing something different about where this is going?
Zhen Lu: I'm sure you're seeing some interesting things from where you are at StrataEdge, but I'd say two things. First, I'm used to being a software developer in a world where compute isn't the limiting factor and doesn't make up most of the cost. Think about SaaS businesses five or ten years ago: the underlying compute didn't cost much compared with the software running on top of it. That is not the regime we're in with AI, which makes it confusing for a lot of developers, because they're not used to operating like this.
Second, because everything is coming so fast, everybody is trying to drink from the fire hose. When you're in that state, you're not very interested in optimizing for cost unless you're a large enterprise that already runs inference at scale. That means a lot of developers, whether they work at Fortune 500 companies or on their own side projects, use the frontier model labs as the easy button. In a lot of cases, they're using a sledgehammer to drive the proverbial nail. Let's be honest, we're all guilty of this.
Over the next several years, we'll see a lot of changes in how software developers operate, because these things will shake out. Compute will ideally get cheaper, but as developers we'll also need to get a lot more efficient at harnessing it. That means writing better software layers, using smaller models for specific use cases, and running fine-tunes. That's a different class of problems from the ones most people have today. A lot of people are just thinking: how many tokens can I get out, what's the best model I can use, and how do I integrate it into my existing business systems? We'll see a rapid evolution of that in the next couple of years, which I'm excited to see.
Allan Carroll: I know developers who have "ultrathink" hard-coded into everything. There's a command they can just smash to get the most tokens and the most thinking on everything they do, even the smallest problems.
Zhen Lu: Most people right now are more interested in getting their problems solved than in what it costs, until they start getting bills that are eye-poppingly large. Actually, I'm curious: what's the largest AI bill you've seen so far, either at StrataEdge or anywhere else?
Allan Carroll: For a client, in the millions of dollars, for sure. I see it as a barbell. Some leaders tell me, "I'm spending that much, but it's 3 percent of the budget for my engineering team. It's in the noise, even though the CFO wants me to start looking at it because it's a real line item." Others say, "I'm spending a lot more on AI than on a bunch of engineers. Are the engineers more effective, or is the AI?" and they're actually cutting back on AI usage because of that.
Zhen Lu: We're definitely seeing unprecedented demand. Demand from our customers is pretty insatiable, and I'd be disingenuous if I said it was all rainbows and unicorns. It's a hard problem to solve.
Allan Carroll: We hear a lot about power as another constraint. Can you even get the power to run all these chips, the compute, and the cooling?
Zhen Lu: That's a whole different story. From where you're standing, are a lot of your clients running on-premises or in other clouds? If they're talking about power and cooling, are they building their own data centers or running in-house?
Allan Carroll: Some are certainly thinking about it. When they go to the hyperscalers and can't get the compute, or can't get enough quota, somebody says, "Just run in-house. Buy your own GPUs, then you control it, and you can bring your costs down because you can run those GPUs all night and get as many tokens as you want." They think about it, but then they run into the same constraints and come back to the decisions we talked about earlier: "I can run an open-weight model and probably get good enough results, but why wouldn't I want the frontier model on everything? What trade-off in intelligence am I making by going on-premises and doing this myself?"
Zhen Lu: I think we're still in the era where everybody can, and probably should, start with a closed-source frontier model. Then there may be a fork in the road where they have to choose whether to continue with that or move to something more open. What will make that decision is the kind of use case, how private your data is, how much control you need, and how mission-critical it is. The more mission-critical it is to your business, and the bigger your data moat, the more I see those use cases trending toward open models over the long term rather than closed frontier-lab models.
Allan Carroll: That will definitely be the big driver. The more interesting conversation we're having with clients is about where the data goes, who gets it, who has access to it, and what models it's used to train. I think the labs are fairly upfront about when and how data is used. But when you have a big data moat, you still want to control your own destiny. One thing we're seeing is that everybody wants to build these new tools inside their own security boundaries.
It's a big change in how software gets shipped and developed. Instead of trusting a SaaS provider that runs on its own hardware and gives you a link to use, they're saying, "It's a lot easier now to build personalized, customized software. Why don't you deploy it inside our security boundary, in our environment, whether it's on AWS or not? Then we can trust it more, because we know where the data sits." And then, of course: which LLM do you use, how do you run it, and how do you handle egress and those protections?
Zhen Lu: One thing I wanted to ask you: how are your customers thinking about their deployments? Are they getting more complicated or simpler over time, if you've got enough data points?
Allan Carroll: In the early work, everyone was taking a shot at the biggest-value problem they could. It was simple: take some off-the-shelf pieces and see how quickly we can solve this problem. To your point earlier, it didn't matter what cost we threw at it. If we could solve it, it was worth it. One executive told me that even if each output cost $5,000 in tokens, it didn't matter, because of what the human effort cost.
That started simple, but once they put it together, they said, "Now it can do these other things." They're also thinking about how to make the agent more proactive, so it runs on your behalf in the background, even speculatively. Can it produce an output? Then you say, "Sure, I'll take this one. It's good enough." Even if I throw nine others away, it still saves me time.
Zhen Lu: That sounds like a pretty good customer to have. I could use an introduction if they're willing to not have a budget for their spend.
Allan Carroll: As you point out, when the cost hits, that may be a different discussion. But I'm sure we can make an introduction.
It's interesting, coming back to where you started. We've been through a couple of regimes. In the mainframe days, the hardware was expensive and the engineers were much cheaper, so you optimized around that constraint. I grew up building software in the phase you described, where the engineer was the most expensive part of the system and the hardware was cheap. So we optimized by getting Aeron chairs and seven monitors for everybody, because that was cheap compared with the human output. Now we may be entering a new regime where that balances out more.
One thing I heard the other day: as open weights get more stable, and frontier models maybe don't have as much of an advantage, do you build the weights into the hardware and get huge cost savings? That would bring the token cost down, and maybe it balances back out. Or are we going to be in a new regime for a long time?
Zhen Lu: We'll definitely get there. I'm not sure if you've heard of ChatJimmy. It's been around for a while, maybe six months to a year. They did exactly what you said. I believe they etched Llama 3 directly onto the silicon, and that silicon could only run Llama 3, or 3.1, or something like that. If you go to the website, it's live and you can try it. It's absurdly fast, something like 15,000 tokens per second, so it feels qualitatively different.
What's holding that back is agility. As you said, we need some maturity and stability before we can start to see those. But I can see a world where, if the difference is really 15,000 tokens per second versus hundreds, and you're a company running business-critical workloads that needs inference on your specific task, it might make sense to pay for something like that, because the cost difference and the efficiency could be that large. Ultimately we need to see how much it costs, because that's how you'll determine the return.
Allan Carroll: It certainly brings latency way down. That's incredible. I can imagine three or four industries where the latency alone is worth billions of dollars. For others, maybe you get economies of scale, where somebody builds the chip and inference costs come way down.
We've talked about this before: where's the bottleneck in the system? For our customers, even at 15,000 tokens per second, you couldn't saturate that with all the context and data you need from their business. Where does all that come from, and how does it get in there? The bottleneck for these groups will consistently be getting data to the LLM fast enough to take advantage of that speed. So it's an inference-cost problem, and they're all thinking more about how to get all of their company context into the right place so they can take advantage of it.
Zhen Lu: It's incredibly difficult. This is why I think there's a world where fine-tuning becomes a lot more commonplace. Even with incredible caching, there's only so much data you'll want to stuff into context on every request. That ends up ballooning your costs, and it adds latency. If you can take some of the more stable knowledge and behavior and fine-tune it into the model itself, you reap longer-term rewards. Again, it comes back to maturity. If you're still developing, it doesn't make sense to make that investment early on. But as things mature, I think that's where we're headed. I'm waiting for the day I can get custom silicon drop-shipped to my house like a T-shirt and run inference on it. That would be very interesting.
Allan Carroll: You mentioned fine-tuning. How does that affect our ability to push the hardware? If it's more efficient and the context is built into the model, maybe you don't have to push it as hard. What kind of breakdown are you seeing between people who fine-tune and people who don't on your platform?
Zhen Lu: The most successful customers we have on Runpod are the ones that fine-tune, run inference, and do research. They're running the full software development lifecycle for AI. They're generally the ones building their businesses on it, because they recognized early that they couldn't build a long-term business moat if they didn't own the model and have a custom model that could do things other companies couldn't. Some businesses build with more of a SaaS mindset, where they wrap the model with systems, context, or user experience to differentiate. That's a fine start, but ultimately I see these businesses needing to customize at the model level to succeed. That's been exciting to see.
There are several reasons to go open source. I call them the three C's. There's cost, which we've talked about. There's control, which we touched on. Recent model changes and government regulation really woke a lot of enterprises up to the fact that if they depend on somebody else's model, it can be taken down at any time or deprecated for whatever reason. That's not great if all of the systems you rely on break because of something you have no control over. I think that's a pretty decent reason for customers with business-critical workloads to run them on open models, on infrastructure they control.
And then there's customizability, which is the fine-tuning piece. What I've heard from enterprise leaders I've spoken with is that they've gotten results better than the frontier models by picking smaller models and fine-tuning them. Whether you call that distillation or something else, we have companies running robust reinforcement-learning loops. That gets them better accuracy, lower cost, and more consistency than a very general model, for use cases that are well scoped within their businesses.
Allan Carroll: I've seen a lot of the same, especially where something almost works with the general model and fine-tuning provides the bump for that specific use case. It becomes a lot faster, because the model already understands the context. It isn't burning through the context window to get to the answer, and it solves problems where the general model doesn't have enough specificity for the places we work.
Allan Carroll: It's interesting to think about the AGI question as people build products like the ones we're building with companies, and then a new model comes out. How do you think about building these stacks when model capability changes every six months? If you point your existing tools at the new model, they work somewhat better, but you have to re-architect around it.
Take the audit work we've been doing. We built a harness, and we've helped companies with financial audits where we improved accuracy, efficiency, and cost, and found errors people weren't finding. But that takes a harness and a lot of work around it. You can imagine that if AGI arrives, you just tell the agent, "Run an audit," and that's all you have to do. Will we get there, and how much does six or twelve months buy you on that path?
Zhen Lu: It's a hard question. Nobody has a crystal ball. I don't subscribe to "let's not do anything and wait for AGI to solve our problems." Even the definitions are ambiguous, but say it's the superintelligence that can do everything, including an audit. To get there, I think we need a lot more data. Even all the data in human history, coding or otherwise, isn't quite enough. I think we're simultaneously farther from that than most people think, but maybe closer than others think. I'd give it at least five to ten years.
In that time we have to figure some things out, which is why I think it's incredibly important for businesses to invest in harnesses. Between now and then there will be a lot of models, and you don't want your company's processes to depend on any one of them. You want them to be fairly pluggable, because models come out monthly, even weekly. You want all the enabling functions around the model: clean data and somewhere to store it, because it's garbage in, garbage out, and tool calling and access to those tools. We haven't touched on governance and security, but that's incredibly scary at the company level. Give a monolithic "god agent" access to everything, and things can go really wrong.
And then testing, whether you call it evaluation or something else. Because these systems aren't deterministic, you need confidence that when you change things, you won't regress into things that don't work. If it's business-critical, you can't afford that. So investing time and effort in the parts that should change less often than the model is very prudent.
Allan Carroll: I certainly agree. The models will change, and we don't know what the future brings. Waiting doesn't make sense when you already have such a powerful tool.
We spend time on these projects in three buckets. The first is figuring out what we should do. The second is doing it. The third is making sure we did the right thing. In software, we'd call that the spec, the coding, and then evaluation or validation, but every one of these workflows has the same pattern. The middle part, doing the thing, is getting so much faster and cheaper that it makes a lot more sense to spend your time and resources on specifying what to do and validating the result.
Zhen Lu: We agree the middle step is getting compressed. Do you have an opinion on whether you should spend more time and effort on the beginning or on the end?
Allan Carroll: That's an interesting thought experiment. I might say they're roughly equivalent. If you can build the evaluation that knows what the right answer is, that's almost the same as writing the specification for what you want, what the right thing to build is, and how to get there. We're much better at validating whether results are right, and the machine is much better at doing the work. When the answer is right, the green light turns on.
Zhen Lu: Now we've somehow gotten into test-driven development. I can see that. The rub is that I don't think anybody has figured out how to test these things effectively. We're getting a lot better. There are new introspection techniques, whether you go to the trace level or run a battery of model evaluations. But it's not a solved problem, and I'm curious to see how it evolves over the next couple of years.
Allan Carroll: I think so too. Even at the business level, most people have an understanding of what they want to get out, but it isn't very precise. Sometimes it's "I'll know it when I see it."
Zhen Lu: How do you think about agents and managing agents in the enterprise? A lot of people anthropomorphize them. We see agents as the analog of people, maybe not very capable people, depending on the agent, and generally not able to operate at the level of a peer. So there's an amount of management involved. People talk about treating agents as teammates you need to manage, and a lot of those activities, like setting expectations and giving them the right specs, are very similar to traditional management.
The major difference, aside from capability, is the timescale they run at. Have you had conversations with your customers about this mismatch, where humans operate in hours and agents operate in milliseconds? Think of it as being a people manager who could only have a one-on-one with each person every three years. How do you keep things from going off the rails? It's an interesting thought experiment when people are trying to drive these things as efficiently as possible without letting them go off the rails.
Allan Carroll: It's a big topic for us, especially in finance, where a human in the loop is a critical piece of governance, compliance, efficiency, and accuracy. We're still having the conversation where they say, "Every time it takes an action, I want somebody to look at it and approve it." But I can't do that if it has effectively run for three years first. There's no way for a human to work on that timescale and approve as fast as the agent could go. What they've come to terms with, at least for now, is that they won't necessarily saturate the agent, because people can't run at that speed.
Zhen Lu: So they're making an intentional decision to slow down because humans must be in the loop. And if humans are in the loop, you may be slowing down by more than a thousand times compared with what the agent could do by itself.
Allan Carroll: Right, but they still get a huge speed-up over what's possible without the agent. For example, it runs a batch of work in the background that takes five minutes, and nobody looks at it until the next day. But somebody else didn't have to do work that might have taken them days.
Zhen Lu: That's fair. I guess they have to make peace with the built-in inefficiency. Either you run it yourself and figure out how to keep that compute busy doing something else, or you ask somebody else's service to multiplex it for you in a possibly multi-tenant environment, which comes with its own security and governance concerns.
Allan Carroll: Which comes back to our earlier conversation. Most of our clients are less concerned about saturating their Anthropic limits than about where their data goes, how it gets there, and who uses it.
Zhen Lu: That's not surprising. Do you see that changing? This gets into accountability. Is there a world where we let AI run with less human involvement? Where does that happen first, and how do people get comfortable with it? Do we need new kinds of insurance to protect against agents running amok and destroying things?
Allan Carroll: Of the insurance companies we work with, I have yet to see one offer "AI run amok" insurance, but to your point, that may be coming before too long.
We're seeing it play out in software, because developers are at the forefront. I don't know if you know any developers who run in "YOLO mode," with no human in the loop. They let the agent do what it's going to do. But often they've set it up in a sandbox. They've protected against the biggest risks by controlling what the agent can access. It's not in full god mode. They don't look at every command it runs, but they do have guardrails. So I think you're right that there will be institutions and guardrails. Understanding those deterministic guardrails is a big piece of work that a lot of people are putting time into right now: how do we let it work without a human in the loop, but within certain boundaries? And coming back to validation, we spend 20 times more on validation than we did before, because we're willing to let it run independently, but then we make sure we get the right output.
Zhen Lu: I was at a roundtable in New York recently where there were two opposing opinions about this, specifically around developers. One camp said that ultimately the human has to take accountability. It doesn't matter whether the developer used AI to write the code; if they're associated with the pull request, they're on the hook. I was surprised that someone brought up the opposite view: if you expect your developers to innovate and not be risk-averse, you can't hold them accountable that way. That applies today and in a future where things are ramping up. As a software engineering craft, we need to find the next way to think about this, because that won't cut it if you want to accelerate.
Zhen Lu: Which camp are you in?
Allan Carroll: I've been changing my mind over the past six months. I started in the camp that developers should take accountability for what they put their name on. It harks back to the old IBM adage: a machine can't make a decision, because a machine can't be held accountable. I think there are certainly parts of that I still hold.
At the same time, I don't hold my developers accountable for the output of their compiler. If the compiler makes a mistake, that's not their mistake, because there's no way they could ever inspect every compiled program. But I still look at the validation: did the output do the right thing in the long run? When we talk with our customers about the help agents are giving their teams, it's the same. The outcome still has to be owned by somebody, even if the small details of what the agent did aren't held to the same accountability, so the team can move faster. In the end, they still want their finances to be right, so someone is accountable for that, whether the agent or a person did the work. And when you're producing software, the system still has to work and do what your business needs.
Zhen Lu: I'm in a similar boat. I was looking for a black-and-white answer, and the comfortable answer was, "As a developer, you have to take responsibility." I think the final answer is more nuanced than that, but I don't have it yet, so you're in good company.
Check the business need, alternatives, data, responsibilities, results, and long-term cost.
Linda Apsley
Managing Partner & CEO
Software gathers a person’s records, puts them in order, and links each entry to its source for a reviewer, with masking, access control, and logging built in.
By Linda Apsley
Buy if a product fits your records and review process with configuration. Otherwise build. Either way, test it on a quarter you have already audited.
By Linda Apsley
Clear limits on what the agent can do, tests on real examples, people reviewing the decisions that need judgment, monitoring after release, and a team that owns it.
By Allan Carroll
Check each bordereau against the delegated authority agreement when it arrives, show the evidence, and send only the discrepancies to auditors.
By Allan Carroll
Usually four to eight weeks, if you test one piece of work, have the data at the start, and agree up front how to judge the result. We built one for BCG in four weeks.
By Allan Carroll
Ask every firm who will do the work, what they have running in production and how they measured it, what it will cost to run, and how your team will take it over.
By Linda Apsley
Verify income, assets, and credit at application, track every condition, start each task when it is ready, and automate routine work. Underwriters keep final approval.
By Allan Carroll
Building, data preparation, model usage, infrastructure, human review, monitoring, maintenance, and training people. Five of these keep costing money while the software runs.
By Linda Apsley
More articles
How we reduced quarterly audit work and what it takes to review, operate, and maintain the software.
By Allan Carroll
Give teams consistent records, clear ownership, and the source history needed to investigate differences.
By Linda Apsley
Define the work, check the data, price ongoing use, and prepare the people responsible for the result.
By Allan Carroll
More videos
Linda Apsley with Vlad Lukic of Boston Consulting Group
15 min
Linda and Vlad Lukic, who leads BCG’s Tech and Digital Advantage practice, discuss why AI projects fail and six questions to answer before funding one: the effect on profit and loss, build or buy, data readiness, governance, when to stop, and lasting advantage.
We can help assess the work, the data, and the full cost of building and operating it.
Talk with us