AI Wars: Managerial Economics Strikes Back

Episode 20: AI Wars - Managerial Economics Strikes Back
===

Morgan VanDerLeest: [00:00:00] Hey everyone, and welcome back to the PDD Podcast. Oh my effing God, what a whirlwind of a year that it's been so far. So Eddie and I both are busy with our respective careers, not least trying to keep up with success metrics for AI adoption, which are redefined basically daily. Techniques, platforms, frameworks, and even more critically, organizational strategies that we're trying to figure out the place of artificial intelligence in businesses and A lot of those have been rising and falling at a pace that makes the JavaScript framework wars of the 2010s feel absolutely quaint.

And back then, people used to joke they take the weekend off and come back to find three more JavaScript frameworks, and today even crazier

Eddie Flaisler: Totally. We've been sitting on today's episode for quite some time, even though listener questions around this topic have been there for at least two years now. We waited not because we didn't have anything to say, but because, as we often mention in our series, we seem to be living [00:01:00] in an era marked not just by insane technological advancements, but also by substantial changes in societal norms and priorities.

We needed to see where this was going, and while neither of us is a futurism expert, we're finally starting to see some convergence worth analyzing.

Morgan VanDerLeest: Okay. I'm gonna read this week's listener question verbatim. It says, "Dear PDD," and then the words, "What the actual fuck?" crossed out, and then, "I am a VP of engineering at a late-stage startup reporting to our CTO.

We're fairly large and well-known by now, so unfortunately that's all I can share. During my time here, I found this to be a very healthy workplace with mature leadership and thoughtful, well-communicated decisions. In many ways, that hasn't changed, but we're now facing a frustrating situation. We need a Series E, and we've been struggling to raise it.

We got on the AI first train relatively early, but apparently even that is no longer enough. Potential investors now seem to want to micromanage exactly how we do [00:02:00] AI. Just this week, a principal at a VC asked us to confirm that we are loop maxing."

Eddie Flaisler: Wait, Morgan, did he say looksmaxing like the bros on TikTok who do cutting?

Morgan VanDerLeest: Eddie, what what are you watching? And no, he said loop maxing

Eddie Flaisler: Oh, and that's better

Morgan VanDerLeest: Now, that's actually a good question. Let me-- Hang on. Let me finish reading this, and then we can talk about loop maxing itself. All right. He writes, " Our C-level leaders are really great, but we need to survive, and they too are starting to cave under the pressure. They're forcing me and the other VPs to make both personnel and technology decisions that we're not comfortable with.

I want to make this better, but I don't know where to start. Send help."

Eddie Flaisler: I just can't get over the fact that TikTok terminology is apparently official language in technology due diligence now

Morgan VanDerLeest: Eddie, we've got some bigger fish to fry. All right, cue the intro. Let's get into it

I am Morgan.

Eddie Flaisler: And I am Eddie.

Morgan VanDerLeest: Eddie was my boss.

Eddie Flaisler: Yes.

Morgan VanDerLeest: And this is PDD: People Driven Development.

Eddie Flaisler: So tell me more about the fish we're frying

Morgan VanDerLeest: [00:03:00] Well, for one, I think this question falls into this broader category that we've handled on this podcast where someone is seeking guidance, but it's not entirely clear what they actually need. Now, the issue here seems to be expectations from venture capital, but what can a leader who's not even on the C-level actually do about that?

Eddie Flaisler: Well, you're definitely not wrong that it's not obvious how a non C-level leader can directly help solve a fundraising challenge or, uh, set expectations with investors. What I do think is that a VP or even a director of engineering for that matter, in an organization of this size, is in a uniquely advantageous position.

They're close enough to the technical work to understand the implications of the decisions being made, while still senior enough to have a seat at the table where the right questions can be asked and proposals can be made.

Morgan VanDerLeest: t Now, as we've mentioned in the past, there's nothing we can do to solve a situation where the board or leadership does not believe in collaborative, mission-driven decision-making.

But for the purpose of this discussion, I [00:04:00] think we can assume that's not the case and that there is room for fresh ideas

Eddie Flaisler: That's exactly right

Morgan VanDerLeest: Okay, so what does this proposal look like?

Eddie Flaisler: Well, as it so happens, a few weeks ago, Eric Ries, the father of The Lean Startup and founder of the Long-Term Stock Exchange, released a book called Incorruptible: Why Good Companies Go Bad and How Great Companies Stay Great.

Morgan, at a time like this, when some of the decisions we're seeing in our industry just feel completely upside down, I needed that book like oxygen. It genuinely restored my faith in people. I highly recommend it. The central idea behind it is that organizations often end up making strange decisions, not because the people involved are incompetent or malicious, but because otherwise good leaders are operating under pressures that gradually distort how decisions get made.

In this book, he introduces tools for helping organizations respond to pressures without creating new problems in the process. What I think we should do is borrow [00:05:00] three of those concepts and adapt them into a framework for the problem we're discussing today: financial gravity, structural integrity, and coherence.

Everything, of course, in the context of AI engineering, which is what our listeners specifically asked about

Morgan VanDerLeest: You know, when we first started turning this into an actual, concrete episode, the name Eric Ries resonated right away, and I hadn't gotten a chance to read "Incorruptible" yet, but I am really excited to dig into how those frameworks can apply to tech industry in general, and particularly in the context of AI engineering.

So really looking forward to that. Now, before we dive in, I think it's worth saying a few words about the title of this episode. It's called "Managerial Economics Strikes Back," and I want to make the connection between our listener's question, the framework from Eric's book, and managerial economics.

Eddie Flaisler: Yeah. So the natural way to do that is to explicitly define what managerial economics is. Managerial economics is the discipline of making decisions under real-world constraints. Traditional economics [00:06:00] often assumes ideal conditions so it can explain or predict specific phenomena. Managerial economics, on the other hand, is about deciding under reality, opportunity costs, scarce resources, risk, and considerations that don't come with an obvious price tag.

You of all people know that while things like employee morale, technical debt, or customer relationship don't come with invoices, they still have real economic value.

Morgan VanDerLeest: I can't tell you how happy it makes me that I get to actually use my degree in economics for this. I love it. Now, I see where you're going with this. So first, let's take a step back. I think it's about educating ourselves on how to decide in nuanced situations.

And second, following Eric's framework and practice requires a constructive approach. Like, just saying no is counterproductive. You have to be able to say, what yes looks like, and that requires actual systematic thinking

Eddie Flaisler: That's right. So let's actually call out the biggest [00:07:00] constraint people are managing through right now

Morgan VanDerLeest: You mean that everyone has just lost it?

Eddie Flaisler: sure. Maybe in a more PG-friendly framing though

Morgan VanDerLeest: Well, I think it might actually be useful to separate symptom from cause here. So for symptoms, I think we're seeing two things happening at the same time. One is that companies are adopting AI much faster than they can actually realize value from it. They roll out copilots. They redesign workflows around agents.

They mandate extensive use of that technology

Eddie Flaisler: # loopmaxing

Morgan VanDerLeest: See, you love that word, don't you? Now, the point is that the way we work and the way context is organized inside of the company has not caught up and sometimes is fundamentally not aligned, so they struggle to capture the actual value that they're expecting. And second I'd say is premature organizational restructuring in preparation for these AI productivity gains that just haven't materialized yet.

You know, most notably, we've seen a wave of layoffs and hiring freezes that companies have linked to AI investments or expected [00:08:00] AI-driven efficiencies

Eddie Flaisler: Right. And then there are the causes, which seem to be the actual constraint. I think there's been quite a bit of speculation as to why all of this is happening, but at least the research I'm exposed to converges on one foundational point. AI disrupts labor, which makes the market reaction to it very different from what it was for previous waves like blockchain or social media.

And because it disrupts labor, it triggers two of the most basic instincts in every business leader. The fear of competitive disadvantage and the preoccupation with marginal cost. Now, the thing is, Morgan, there is nothing wrong about an executive being concerned with these. I would in fact be worried if my C-level wasn't actively thinking through how to differentiate and how to make sure operations are sustainable.

The issue is that people thinking through these challenges are subjected to what Eric calls financial gravity. Essentially, the pressure to prioritize the success of a specific future transaction, be it a funding [00:09:00] round or an IPO or next quarter's numbers over the things that actually make the organization successful in its mission.

Accumulated assets like loyal customers or motivated talent or earned trust, and of course, the innovation they make possible are a time tested recipe for extraordinary long term value, both financial and mission related. The problem is that their benefits are realized over years and are difficult to attribute to a specific set of decisions.

Whereas layoffs improve margins this quarter, AI announcements impress investors today. Cost reductions are visible immediately

Morgan VanDerLeest: Yeah, I can see that. I think the layoffs are a particularly good example of it. Savings are immediate, margins improve immediately, and investors and stakeholders understand the story without requiring deep understanding of the domain . The long game is harder to paint a picture of. Combine that with the fact that individuals making these decisions often experience their company's production AI primarily through, [00:10:00] quote-unquote, "happy path demos," and by virtue of their role, are just not always close enough to the work to tell whether what they're seeing is truly a net positive outcome or a marginal value purchased at a disproportionate cost, and you're left with very little to discourage that pressure from taking over

Eddie Flaisler: Absolutely. So now we can finally summarize. The managerial constraint of the AI era is that maximizing the next transaction and maximizing AI's long-term value often require opposite behaviors. Organizations realize value from AI when it has access to appropriate context, when it can leverage institutional knowledge, when it is managed through thought-out feedback loops, and is directed by strong human judgment.

Financial gravity, by definition, pushes organizations towards reducing costs and producing immediate results. And unfortunately, when you do that, these very capabilities are often the first [00:11:00] to go

Morgan VanDerLeest: Well, holy shit

Eddie Flaisler: I know, right? Tragic

Morgan VanDerLeest: All right. Well we are here after all to offer some consult, so let's start working through the framework that you borrowed from Eric Ries. Now, you mentioned step one, financial gravity. We already defined what that is. Now what do we do about it? Where do we start?

Eddie Flaisler: Everywhere. So in Eric's book, he talks about recognizing financial gravity and being mindful of it when discussing decisions. We're going to build on that a little. You don't need to be a seasoned leader to know that when humans are dealing with something high stakes, they're not always open to listening, let alone being told they might be susceptible to pressure, present company included.

So to me, the practical manifestation of that isn't saying, "Let's acknowledge we're struggling." It's making the numbers and trade-offs as visible and legible as possible to the people making the decisions. And to do that, I think we need a very short review of AI mathematics

Morgan VanDerLeest: Go for it

Eddie Flaisler: Rule number one: Can we please stop [00:12:00] counting tokens? It's the AI equivalent of starting a mortgage conversation with, "Here's how much I can afford every month." Your broker just smiled and started shopping for a second boat

Morgan VanDerLeest: seriously, what am I gonna do if I'm not just gonna sit there and watch my token number approach budget? That's what I do on the daily, Eddie, come on

Eddie Flaisler: Well, there are three pieces to this. Hear me out. Bad price signal, bad value signal, and bad planning signal. First, cost per token or million tokens or whatever people use when choosing models is a quoted, not measured, but quoted by the provider number.

It reflects costs in a steady, predictable request volume where the hardware is fully utilized and requests can be batched efficiently. Nothing like the spiky variable traffic of a real product. Second, token spend is somehow used as proxy for AI productivity. More tokens apparently means you're doing more with AI, but it can also mean you generated a long document which nobody read, or asked the [00:13:00] same question multiple times, or wrote volumes of throwaway code because the guidance wasn't specific enough.

Third, and that's probably the most scary, in an agentic workflow, it's practically impossible to predict token spend because you don't know how many steps the agent will take to solve the problem before it does.

Think of a naive react loop. Cumulative input tokens conception grows roughly quadratically with chain length, assuming tool outputs are roughly uniform in size. A ten-step workflow where each step adds five hundred tokens of tool output creates about twenty-seven thousand tokens of cumulative input processing.

And that's before counting the system prompt or the final answer generation. So the final transcript is only five thousand tokens long, but the model reprocesses that growing context window from scratch on every call.

You're not budgeting for a task, you're budgeting for a number of steps you won't know until the bill arrives

Morgan VanDerLeest: That sounds like a great billing model. Why aren't we all doing this?

Eddie Flaisler: Oh, [00:14:00] totally. I asked myself the same question

Morgan VanDerLeest: All right, so jokes aside, all great points. But if token spend is not the number we want to surface, then what is?

Eddie Flaisler: Well, the two considerations that still make sense to everyone, even when we're panicking, are one, value analysis. So does our AI, and the operative word here is our, not some theoretical war story on LinkedIn, our AI deployment actually create value? If so, what kind of value and to what extent? If not, why not?

Morgan VanDerLeest: And I think this is where it's important to distinguish between internal facing AI, and customer facing AI. So measuring value to customers has its own classic metrics that still work well here. NPS, retention versus churn, adoption rates, revenue expansion, and so on.

Internally is where things become more challenging because traditionally, technology organizations haven't been particularly disciplined about measuring work So the foundational work to understand AI value internally may actually be improving how [00:15:00] we measure engineering. Now, using metrics like person hours saved, time spent building, and idea to delivery time help answer whether we're doing the same work faster.

But there are also other really interesting things to measure

Whether people who previously couldn't contribute are now able to create work that actually ships, whether we're finally getting to work items that sat in the backlog for years because it wasn't economically viable before, and whether more experiments are actually making it all the way to production.

And when I say production, I mean production ready. You didn't just push whatever. So the questions for internal facing AI are essentially, are we actually saving time? Did we lower the barrier for contribution across the company? And perhaps most importantly, are we now able to do things that we simply just couldn't before?

Eddie Flaisler: I find this very astute, especially the observation that investing in how we measure engineering is now more important than ever, because we're suddenly making a whole new class of decisions based on this data, and some of those decisions are extremely expensive. That's what makes [00:16:00] things like PR hygiene or ticket management so important across all stages, even though so many love to dismiss them as red tape.

Clean Jira, clean GitHub data, it gives you an unbiased, objective benchmark that you can actually compare against. Once you have that baseline, measuring AI effectiveness becomes much more tractable

Morgan VanDerLeest: Absolutely. Now, what's the next consideration?

Eddie Flaisler: The second one is cost analysis, and this is where it gets interesting

Morgan VanDerLeest: Because you said not to count tokens

Eddie Flaisler: Right. So here's the paradigm shift. So far, we've been talking about financial gravity from the angle of how do I implicitly tell the business this thing may not be cost effective? But the thing is, Morgan, with all due respect to pushback being part of the engineering leader's job, we were hired first and foremost to bring solutions, not problems.

So the first numbers to look at are not the numbers you're showing business leadership. They're the numbers you're asking your own team to show you. Efficiency [00:17:00] metrics. In the world of AI, this turns out to be a surprisingly small set of metrics. Cache hit rate, request distribution between different model classes, quality delta after quantization for self-hosted models, and cost scaling behavior as usage grows

Morgan VanDerLeest: All right. Whoa, Eddie, hang on a second. Quantization, caching, this sounds very solution specific

Eddie Flaisler: Only it isn't. People can adjust the terminology as needed, but ultimately, we're always trying to answer the same set of questions. Did we minimize calls to the model to only what was actually necessary? Did we choose appropriate models for specific tasks, which also means not using a flagship reasoning model to summarize a two-sentence paragraph?

Did we apply the industry's current best-known optimizations to reduce cost without materially sacrificing quality? And did we design context retrieval for our agentic loops such that they don't have to reread the entire history at every [00:18:00] step?

Morgan VanDerLeest: By the way, I love that model routing is suddenly a thing, as if it required some scientific breakthrough to realize that you probably shouldn't route every AI task to the same flagship model. I think it speaks volumes about the historical lack of cost discipline that we've talked about.

We ended up with some of the best engineering talent in the world. Building systems for efficiency simply wasn't a design constraint, because for a long time, cost simply wasn't a design constraint

Eddie Flaisler: Amen to that

Morgan VanDerLeest: Okay, so we have the efficiency metrics in place and you as a leader know your team is doing everything it can to use AI efficiently and predictably in your own systems. Now what?

Eddie Flaisler: Now we go back to the number we surfaced to business leadership, and that's cost per useful output. I'll say it again, cost per useful output. Here, the math gets a little less straightforward, but totally worth it. The first step is defining what useful output means for you. For customer-facing AI, it can be a completed task or an accepted suggestion or a resolved [00:19:00] support ticket, and so on.

For internal-facing AI, it can be a merged PR, an incident resolved, a migration completed. Again, these are all things that require getting into the habit of tracking work properly so you have something objective to measure against.

Morgan VanDerLeest: Absolutely

Eddie Flaisler: So now you have your useful outputs tracked. The next step is looking at your total AI costs, and this one gets particularly tricky

Morgan VanDerLeest: Because this is where you can disappear down a philosophical rabbit hole. What's included? Human review costs? Dedicated infrastructure? Observability stacks? Failure costs?

Eddie Flaisler: That's exactly right, and this is why I say every organization needs to decide what it's going to treat as cost. As long as there's internal alignment, that's all that matters. I think the typical components here are AI vendor spend, GPU and inference costs if you're self-hosting, and AI specific tooling licenses

Morgan VanDerLeest: Makes sense

Eddie Flaisler: And that's basically it, because cost per useful [00:20:00] output is total AI spend divided by the number of useful outputs. The more interesting question is what you actually learn from movements in this metric. And the honest answer is, in a vacuum, nothing. But when discussed properly, it forces you to ask follow-up questions, which are actually the important part

Morgan VanDerLeest: So for example, cost per useful output can go down because you improved efficiency or because success rates improved and fewer attempts are now required to get to a good outcome. But it can also go down because you transitioned to a cheaper model, which might be perfectly fine, but only if the quality numbers are still looking good

Eddie Flaisler: That's right. And it can also mean users stopped attempting complicated tasks because we suck and now only use AI for trivial stuff. So the metric improved, but the value actually dropped.

Morgan VanDerLeest: And similarly, cost per useful output can increase because efficiency deteriorated or quality dropped, so you need more attempts to achieve the same outcome. But it can also increase because adoption is [00:21:00] growing faster than optimization, and we're consciously accepting that because, say, the data generated by this activity helps us build future capabilities

Eddie Flaisler: And I think cost per useful output is a perfect example of a metric that can drive an enormous amount of great work and positive outcomes for both customers and the business. But if you allow surrogation, as Eric Ries calls it, to creep in, meaning the thing we created to measure success becomes the thing we optimize instead of success itself, then you can easily end up with someone saying, "Everybody, reduce cost per useful outputs by twenty percent," and nothing else.

And you just know that's not going to end well.

Morgan VanDerLeest: And doesn't that sound familiar? Making a metric the target really borks the metric. Metrics are the outcome of good practices. They are not the target. The other thing this brings to mind for me is how does cost per useful output compare to pre-AI? I'm sure there are circumstances where AI spend is more economical than another hire and everything that entails, but it's a very tech industry thing [00:22:00] for us to make up a new metric because it wasn't possible to do before, even though it absolutely was, but we just didn't track it.

I'd love to see R&D organization costs per useful output even just a year ago versus today. But I digress

Eddie Flaisler: No, this is actually a very good point. I never thought about it. This was definitely a thing pre-AI as well.

Morgan VanDerLeest: I feel like we're all building the plane while we're flying it. And with that, did we conclude our discussion of the numbers you surface to help navigate financial gravity?

Eddie Flaisler: Yes, value analysis and cost analysis. That's it. There are plenty of other important conversations to have, like technology dynamics and human team health, but we'll get to those later. In moments of urgency, people rarely have the bandwidth for anything beyond these first two

Morgan VanDerLeest: Makes sense. The shorter you can keep the list, the more memorable it is. Now let's talk structural integrity

Eddie Flaisler: Right. So the idea is basically this. We communicated as effectively as we could. We surfaced the numbers, we did our best, Then our business leadership made top-level decisions with the best intentions and a great deal of [00:23:00] thoughtfulness, but still with plenty of potential for negative consequences if we're not careful about the how.

The question now becomes: What do we need to keep in mind as we execute those decisions?

Morgan VanDerLeest: So when we say negative consequences, what aspects are we talking about? I'm gonna guess something like operational integrity, team integrity maybe something like survivability.

Eddie Flaisler: Meaning?

Morgan VanDerLeest: Well, operational integrity essentially means making sure we improve or at the very least maintain the quality, the reliability, and customer experience of the product or service we're providing as we introduce these changes

Eddie Flaisler: Makes sense. Team integrity?

Morgan VanDerLeest: Team integrity is making sure we don't sacrifice the health or effectiveness of the people doing the work in that process.

Eddie Flaisler: 100%. And survivability?

Morgan VanDerLeest: Survivability is making sure we don't create dependencies, lose capabilities, or just make decisions that limit our ability to adapt as technology and business conditions change. And change they will, as we're seeing right now

Eddie Flaisler: I think it's a very good framework for thinking about [00:24:00] structural integrity, Morgan. And personally, I would like to start with survivability because interestingly enough, I feel like this aspect is receiving the least attention right now, even among critics of the AI rush

Morgan VanDerLeest: Let's do that

Eddie Flaisler: Well, to me, the most foundational aspect of survivability, even before you start debating things like technical direction or cost sustainability, is human succession management. It might sound like it belongs under team integrity, but in fact, it's a business strategy issue

Morgan VanDerLeest: I can see that

Eddie Flaisler: Right? Think of it this way. One of the core responsibilities of leadership is succession management. You never want absolute dependency on a single engineer. Why? Because people leave, people burn out, or their performance changes over time.

The irony is that agents create the same dependency pattern but at an organizational level. Agents don't experience the same work pains human engineers do, so they're not naturally incentivized to write [00:25:00] short, coherent, maintainable code. If agents do most of the work, you gradually accumulate code bases that are legible almost exclusively to AI because they become overwhelming for humans to reason about.

Meanwhile, your engineers get weaker because they've spent less time actually engineering. So sooner or later, you end up needing increasingly powerful and expensive models just to maintain your own systems. Smaller models have already been shown to perform extremely well in clean, maintainable code bases and very poorly in code bases that are difficult for humans to reason about.

So you slowly create an organizational addiction to high-end AI solutions simply to sustain engineering velocity.

Morgan VanDerLeest: Side note, I liked the comparison there of small models understanding cleaner code bases better, and I wonder if there are some guardrails we can put in place for having those more efficient models versus the higher scale models reviewing code and deciding did this get overly complicated because my [00:26:00] smaller model is no longer able to understand it in the same way, and do some corrections that way.

Eddie Flaisler: Well, I actually think it's a brilliant idea, but here's what it ties to: technical documentation was always a human weakness. That's why it's so funny to me when people now say, "What do you mean? You can give AI complicated tasks, just write a spec."

Oh, really? We finally discovered technical documentation. But the issue was we were never pretty good at generating that. If we are able to crisply write a set of instructions for what constitutes good code and bad code, then yes, the model is able to judge based on that. Well, theoretically it is non-deterministic after all.

But in general, it can be. But this is another example of humans having to maintain engineering skills to manage their own AI better, so the AI can manage itself.

Morgan VanDerLeest: Yeah, absolutely. And to build on that, [00:27:00] that's not just bad successor management, it's also consciously stepping into vendor lock-in. The industry spent decades developing patterns to avoid vendor lock-in for economic reasons, and now over-delegation to agentic AIs inherently drives the opposite outcome

Eddie Flaisler: That's actually why I argue it may be time to move back towards a waterfall style model just with much shorter cycles.

Morgan VanDerLeest: Wait, what? Waterfall? I think I lost you completely

Eddie Flaisler: Well, hear me out. The Agile Manifesto did something very powerful. It reframed the basic unit of work around working software, right? The whole picture with the skateboard and the bicycle, then the car. That shortened timelines. Great. It reduced procedural complexity, and it produced a lot of great outcomes.

But it also had a side effect. Activities like design, testing, and documentation, and all the other supporting work became flattened into the delivery process. From the outside, it's essentially a black box. The expectations became that the engineer would [00:28:00] figure it out

Morgan VanDerLeest: Right

Eddie Flaisler: But now you come to the engineer and you say, "Use AI to deliver the black box." So they do just that. The issue is that if you ask it to, and most people honestly do because timelines have become so insanely compressed, an AI model can generate all of the above from a single prompt. It might not be what you needed, it might be really bad, but it's still a complete output.

And along the way, the engineer has very little understanding of what was actually done or how all the moving pieces fit together. And of course, the argument that they should obviously review the code is mostly useless because no human can realistically review a gigantic chunk of generated code all at once and then truly understand what's going on.

Morgan VanDerLeest: I like where this is going, and I don't wanna spoil the outcome, but I'm seeing this play out in real time where instead of a one-shot PRD, ADD, PR, we're seeing, I believe, the best integration of AI, where we can tackle like each chunk of the process [00:29:00] faster, but still have human judgment throughout the whole thing

Eddie Flaisler: That's exactly right. So what I'm saying is this. Since AI assistance can now accelerate every step, we can afford to break deliverables down into structured stages again. That way, we can make sure the outcome is actually what we need it to be, and that doesn't mean going back to a world where a checkbox takes nine months to ship because of all the red tape.

That's how I operated With my engineers in recent years, and it truly felt like the best of all worlds. There were explicit stages with clear deliverables, design reviews, tests, demos, handoffs.

At each stage, the engineer led the discussion and remained accountable for the output. AI can be an incredible teacher and force multiplier. And if you become good at specifying intent, it can generate most of the code and documentation for you and do a great job. But if you want to avoid entropy, engineers have to be allowed to truly own the work.

And when I say own, I don't mean it from a [00:30:00] blame-shifting perspective. I mean it in the sense that you can't hold someone accountable for an outcome if they didn't have the autonomy to decide how to achieve it

Morgan VanDerLeest: This probably falls more under team integrity, but I think the accountability aspect would especially resonate with people. A very common question today is who should be held accountable for AI-initiated changes, especially when they cause an issue? And I think this goes back to the mandate level concept we discussed in the past.

If engineers have autonomy over the work, they're responsible for the outcome. But if they do not, then responsibility explicitly belongs to the person who made that decision for them, period .

Eddie Flaisler: Absolutely

Morgan VanDerLeest: I think we've covered a lot of what survivability means, and I don't think this is the right place for a broader discussion about cost sustainability. But before we move on to the other types of integrity, something should probably be said about the risks of chasing the cutting edge during a time when everything is changing so quickly.

I feel like over the past year especially, I've heard at least once a week that X is dead and everyone should be doing Y instead. And [00:31:00] organizations do. Of course, this has operational implications as well, but I think it's draining and depleting for an organization on so many levels, which makes it very much a survivability concern

Eddie Flaisler: I don't disagree, But I think this issue runs much deeper than the current reaction to artificial intelligence. I don't think there's a single person who has managed in tech and hasn't encountered a product or business partner pushing for what felt like unreasonably frequent directional changes.

They use terms like nimble or customer obsession,

Morgan VanDerLeest: Fail fast

Eddie Flaisler: fail fast, when in reality it's often a combination of a deep sense of scarcity and probably insufficient discipline around informed decision making as well. Now here's the thing. First, it's not our job to therapize or educate people, even though we're literally doing it right now.

Second, you don't want the other extreme either. You want to experiment. You don't want to be driven by fear and stop moving, and you want to stay competitive. But the question is this, both in terms [00:32:00] of decision making mechanisms and technical guardrails and kill switches, do we actually have the, quote unquote, "infrastructure" to take informed bets and pivot with minimal friction if they don't pan out?

This is true across company sizes and stages, and it goes back to what we discussed in the innovation episode. I don't care how pressured you are, if the team isn't spending some time making sure the organization is adaptable in every sense, you're doing it wrong

Morgan VanDerLeest: And I think it's important to call out here that that flexibility is not just somebody up top said we're doing things differently now and mandating that down. It is the actual flexibility of the team to adapt and deliver value with change

Eddie Flaisler: Exactly right

Morgan VanDerLeest: All right. I think we can safely move to operational integrity. This one feels like a whole universe to cover. So I wonder if we should limit ourselves to the aspects that are more foundational than others.

Eddie Flaisler: Yes, I would say we stick to organizational design and production safety. Two very [00:33:00] different concepts, but probably the biggest pillars when thinking about managing through this new era

Morgan VanDerLeest: That actually makes a lot of sense. I think AI introduced a bunch of new failure modes we just didn't have to deal with previously. They're categorically different from traditional software, mostly because traditional software tends to fail loudly. You have Exceptions, alerts, you have spikes in error rates. AI can fail quietly and confidently, and that changes the governance posture entirely. You bias toward assuming wrongness until reliability has been demonstrated

Eddie Flaisler: Any specific modes you'd call out?

Morgan VanDerLeest: There are five I can think of. There's calibration failure, so the model presents low confidence answers with high confidence, making it difficult to know when additional verification is needed. Then there's confabulation. This is when the model fills knowledge gaps with invented but believable detail, and it doesn't say, "I don't know." Obviously, there's nondeterminism, and that's where the same prompt can produce different outputs on different runs, and so you can't [00:34:00] rely on reproducibility the same way you do with deterministic systems. And there's scope creep

Eddie Flaisler: Oh, that's so true

Morgan VanDerLeest: Right? Now it solves the problem it inferred, but not necessarily the one you specified. And then there's the one that's more sociotechnical, but it is an absolute killer. Individual speed without organizational velocity. Agents make engineers faster, which means PRs flood in.

But DORA metrics don't move because the review and deployment pipeline didn't scale with the increase in output volume, and that failure doesn't show up in the agent's output at all

Eddie Flaisler: I feel like your last one is the hallmark of And again, it touches on the fact that AI mostly amplified problems that were already there. If developer productivity was a P zero before, then agent productivity is P minus one. Tackling the other failure modes is ultimately a choice between two models you want to operate in as an organization: augmentation and automation

Morgan VanDerLeest: So by augmentation you mean engineers remaining in the loop [00:35:00] and AI helps them do the work and by automation, I'm assuming you mean the AI system performs the work end to end without requiring human intervention?

Eddie Flaisler: That's right

Morgan VanDerLeest: And how do we make that decision?

Eddie Flaisler: Well, I think the starting point is accepting a permanent constraint. LLM agents are non-deterministic by nature. The same prompt can produce different outputs on different runs. So the question isn't, can I trust this agent?

It's how much discipline have I built around it? For me, autonomy gets calibrated on two axes: how reversible the mistake is, and how mature the, quote-unquote, "harness" around the agent is. The decision on how much autonomy to give the agents has to be a function of these two. Reversibility matters because mistakes are inevitable.

So the question isn't whether the agent will eventually get something wrong, but how costly the mistake becomes when it does. If an agent writes a PR but an engineer reviews and merges it, I need much less evidence that the change is safe than I would with an unattended [00:36:00] automation. One mistake is easy to catch and reverse.

The other might already be in production by the time anyone notices

Morgan VanDerLeest: That makes sense, but agent harness is a bit of a broad term. What do you mean by that exactly?

Eddie Flaisler: Well, the harness is everything around the model that makes its behavior predictable and safe. Feed forward controls like documented coding conventions and architecture rules, you know, like what we were talking about, the small model learning to vet whether this is sustainable or not. We also have feedback sensors like linters and test runners.

You have bounded execution, like limiting the paths the agent is allowed to access. You have isolated environments like sandboxes to contain the blast radius if it goes crazy and tries to, say, delete all of your production data. It has falsifiers, which are basically predefined contracts to check whether you broke something functionally or non-functionally.

Like if latency exceeds one point five x baseline, stop. You know, all these pesky little checks that should have already [00:37:00] been there, but nobody has time for that

Morgan VanDerLeest: I think the agent harness feels like the area that is like most important for technology organizations right now who are really leaning into AI first processes and methodologies because I think there are a lot of other things that are just baked into the newer models coming out. Like you get a lot of wins for free. But the harnesses and the way that fit your organization are things that you have to define for yourself and for your organization. And so whether you have those or not, or whether they are good or not is really differentiating right now.

So do you have a good real world example of what a mature harness looks like?

Eddie Flaisler: I think so. At least from what I've read on the outside, Stripe's in-house coding agents, Minions, are a good real-world example. The surface behavior, as I understand it, is that an engineer tags a Slack bot, describes a task, and, you know, by the time they're back at their desk, there's a PR ready for review, passing CI with no human code in it.

They reportedly have over a thousand PRs like that merged every week. Now, you might wonder what's [00:38:00] special about this because you could get a PR out with cloud code today. But that's not what's interesting. What's interesting is what makes this pipeline reliable on a Ruby code base that moves over a trillion dollars a year.

For one, generic agents don't know Stripe's internal libraries. They don't know compliance constraints. They don't know conventions. So Stripe built Toolshed. It's essentially an internal MCP server with, I think, over four hundred tools. It gives Minions access to internal docs, to code intelligence, ticket systems, build statuses.

Basically, they have access to the same context a Stripe engineer would have. They also built isolated cloud environments for them, essentially dev boxes that spin up in ten seconds with Stripe's full code preloaded, with no access to production or the internet. So the blast radius of a mistake is heavily contained.

Lastly, the agentic loop itself is a fork of an open source agent, but it's interweaved with deterministic steps that cannot be skipped for [00:39:00] the most part. Linters, test runners. So the autonomy line they drew is very precise. Full creative autonomy within the sandbox, zero deployment autonomy. Humans review every PR before merge. The agent can't skip that gate. And critically, Stripe didn't start with a thousand PRs per week. Trust was earned incrementally as the harness proved itself. Discipline, discipline, discipline.

Morgan VanDerLeest: And that's a thing that feels particularly difficult right now, I'd say from within an organization, is there's this push to just make the big changes fast. But like a lot of us have said for years, and as we've said on this podcast, fast is slow and slow is fast .

You can't get to the point where things are, reliable, dependable, if you're not taking the small steps in the meantime. You gotta do small steps first and get to the point where you can take those larger steps more confidently.

Eddie Flaisler: 100%.

Morgan VanDerLeest: And with that, I think we can move to organizational design. We've already [00:40:00] covered a lot of ground today, and so to keep ourselves focused let's see if we can offer one useful principle for organizational design right now and then one for hiring.

Eddie Flaisler: So let me tell you, if there's one thing the work on Mission Control truly drilled down for me, it's that the biggest unsolved problem in tech remains communicating requirements.

And interestingly enough, AI agents have actually made that problem bigger. Think of your options. If it's a document or a Jira ticket used to communicate what you want, then either it's this wall of text and people give up on consuming all of it, or the author isn't particularly strong at technical writing, which is an incredibly common problem, so the message doesn't get conveyed either way.

If it's a mock, then you're missing the behavioral component. It's really difficult to simulate all the possible things a user can do. Eventually, things become slow and brittle in tools like Figma, which I love, by the way, but the limitation is what it is. So now people say, "Fine, let's have the product team vibe code a shell of what this should look like."

[00:41:00] Okay, but then how do you discover all the hidden options inside the shell? How do you know what to explore? So long story short, telling people what to do and having them do what you meant was always a bottleneck. And now with agents that inherently understand the letter of the guidance rather than the spirit, the chances of divergence are much, much higher.

So to me, the only effective organizational solution when engineers become agent managers is making sure that anyone working on something has the full picture of what they're trying to accomplish across the stack. The UX, the customer outcomes, everything.

And you can really only do that in some sort of pod or advisory model, similar to the matrix model Spotify popularized, which we discussed in our reorg episode. Small teams of three to five people working on something they can fully understand, and as such, vet the work of the AI.

And you can also have embeds from specialized [00:42:00] groups like security or SRE whenever deeper expertise is needed. In terms of hiring, I can tell you this. The highest performing engineers I've seen since AI-assisted development became a thing had one thing in common, genuine interest in understanding and solving the problem end to end, as opposed to blindly following the letter of the ticket. And I'm not talking about scope creep or solving other people's problems.

Within the scope of what they were building, I'm talking about user empathy, curiosity about the broader system they work in, things like that. When they built something, they naturally wondered what it meant for the user. That, combined with enough experience to understand design pitfalls and trade-offs, for which, granted, you do need some seniors around, is what I think matters today.

That's why I don't spend time on all the hype around, "I invented a question AI cannot solve, and only humans can." I give people a prompt, and the prompt says, " Implement mission control. You may use AI." That's it. What does it mean for quality, [00:43:00] security, UX? won't tell them that if they don't think to ask first. What's yours?

Morgan VanDerLeest: For me, the idea of organizational design comes back to organization of people design, just like we talk about on this podcast. But it's about how does the organization of people get designed to do work better? Now, we're in a state where, yes, can write code.

It is cheaper and faster to write code than it has been in the past. That doesn't mean that your entire process is gonna get sped up by 95% or whatever these numbers they're coming up with these days. Turning a six-month timeline into two days is something I've heard, which is just wild. But they didn't actually do that, and that wasn't from just writing code faster.

That was about rethinking how processes happen within an organization. And the best way to do that is by talking to the people that are part of that organization and getting not only their buy-in, because engineers see that AI is writing the code now. That's a thing.

There are some people that maybe still won't believe it or still don't want to [00:44:00] believe it, and that's another part of a conversation. But in general, most folks are like, " My job has become, guiding AI or managing AI." It's less about, them doing things themselves. It's important for them to see. So now it's a lot less about getting them to buy into that idea and more about how do they think our organization and our processes can change for the better. As an example, couple weeks back, somebody on the team had an idea, and they said, "Hey, we're gonna experiment with, a totally different process for getting this project done.

We're gonna block an afternoon off. We're gonna get engineering, design, product, and the stakeholders in the room together, and we're gonna go through what's the problem, what are some guardrails for the solution.

Let's build it, get feedback, iterate, and try to launch something by the end of the day." This wasn't a top-down decision. This was a bottom-up. Somebody's like, "Hey, I think we can change how our process works and give that a go." We experimented with it. What did we learn? Didn't happen in an afternoon, but we took what normally could've been a multi-week process and got it done in an hour and a half.

We took [00:45:00] what could've been a ton of back and forth for code and product review, and we started on one Monday, launched by Tuesday the next week with on and off review between various folks.

Eddie Flaisler: I love that

Morgan VanDerLeest: That was huge. This is a thing that we would've had to totally disrupt our roadmap to do before, but somebody said, "Hey, let's try this new process." Now, we retroed on the process. Ton of learnings. We wanna change how parts of that are done, maybe break up some of the pieces, have different points to make people more collaborative through the process.

But What this illustrates for me is for decades, engineers have been some of the best paid people within organizations, not because they write code well, but because they have a particular, expertise, skill set, and acumen that is useful for solving business problems. And that's still true. Just because the code aspect has changed doesn't mean that there's not still a unique problem and an interesting thing to solve, and it's helping folks feel empowered and have that ownership over how can we do this thing better?

And so I think if I had to bubble that down [00:46:00] into organizational design principle, it's keep trusting your people. Your people will have the answer for you, just like it's been the case for decades in other industries. When you were trying to solve something on your manufacturing floor, you didn't ask the people that were managing those people.

You asked the people on the floor who were seeing problems with the machines. It's the same thing today. When it comes to hiring, I think it's important to remember that AI amplifies processes and it amplifies expertise. You still want to check and make sure that those things are there.

All of my engineers now are people that a year ago weren't really using much AI. They're using some, you know, auto-complete and things, but it wasn't nowhere near the same level that we are today, which means that they all learned how to do it in the last year. So I'm way less worried about how well somebody uses AI now, 'cause just because their previous job didn't offer them the opportunity or they didn't work through that well. That's part of my job is figuring out how to help make sure they do that.

But how can I make sure that they have the expertise and the ability to communicate [00:47:00] well? And it's not just chummy in a conversation, but I'm understanding what you're saying, and I'm able to communicate that back out in a way that makes sense to other people. Cause that applies to other engineers, other stakeholders in your organization, even the robots themselves, because if there's anything that matters in that sense, it's the ability to clearly communicate what you're looking for so that you can get a clearer understanding of the outcomes that you'll receive after.

Eddie Flaisler: Such an underrated skill set, especially today

Morgan VanDerLeest: Absolutely Now I believe team integrity is this final piece here before we close for the day with coherence.

Eddie Flaisler: Yeah. So look, I actually had to think a lot about this section in particular, even though from a theoretical perspective, it's probably the one that changed the least. At the end of the day, it's still about managing people. The reason I had to think so much is because, quite frankly, as an engineering leader trying to make it in this industry like so many others, I don't want to say something publicly that would make organizations reluctant to engage with me.

But having thought it through, I do think some [00:48:00] things need to be said out loud. So here goes. As a leader, an employee, and a human being, I have been so incredibly disappointed and heartbroken by the conduct of some leading companies in our sector. Be it the mass Hunger Games-style layoffs, be it forcing people to delegate their entire job to AI with little regard for what it means for their spirit, their career, or their mental state.

The rush to replace humans with AI long before it has been demonstrated that AI can actually do their work merely at the prospect of no longer having to pay these, quote, entitled employees.

I've done a tremendous amount of research around this over the past year, Morgan, and there is no excuse. This isn't about business priorities. This isn't about capitalism, and it isn't about we're not a family, as Patty McCord famously said. This is about fundamental values. Actions we've seen over the past year cannot be easily walked back from.

There are decisions that communicate very clearly that people don't matter, regardless of how much highly [00:49:00] crafted corporate language is used to justify them. I am firmly convinced that if a leader truly cares about the individuals employed by their organization, most of the theoretical frameworks we're about to discuss become intuitive.

So if listeners leave today remembering just one thing about team integrity, let it be this. Remember that you're dealing with human beings.

Morgan VanDerLeest: Absolutely. I could not agree I think along those lines

We mentioned earlier today, and would have to assume that anybody who's really been diving into building with AI. You know there's a wall that you hit that the bot is not able to get past and that you need a human for. And so if you are optimizing for a world where you only have the bot, you know you are going to hit a wall and a ceiling at some point.

I think pretending otherwise is negligent. So as an industry, as leaders, push for the change, see how far you can get, but not to the point where you've [00:50:00] exhausted all of the human expertise that you've built over the years because you're gonna get the immediate "win" because, oh, look at us, we've cut costs, we're all AI.

Six months, a year, couple years down the road, the value that you are actually building for is gonna disappear because you've artificially constrained yourself by removing the one thing that actually made value for your business, and that was the people in your business.

Eddie Flaisler: 100%. And quite frankly, I don't think anyone needs to wait for a few years, Morgan. People are already starting to see this now

Morgan VanDerLeest: It is a little terrifying how fast a lot of this is happening now. That said, there's been quite a bit of research around team dynamics and the human aspect of managing through the AI boom, most notably by Google Cloud's DORA, Atlassian's DX, the engineering intelligence platform Quotient, the MIT Center for Information Systems Research, and ICONIQ Capital.

Eddie Flaisler: Absolutely. Seminal work from some very deserving organizations

Morgan VanDerLeest: Now, do you see their conclusions as [00:51:00] similar or different?

Eddie Flaisler: Remarkably similar, actually. Somewhat different terminology, different audiences, but they all seem to be converging on the same few ideas. And in reviewing these, I actually found something that was both shocking and, in a way, entirely expected. Even though nobody used the term explicitly, the biggest realization was that the Jidoka philosophy from the Toyota Production System, which we discussed in our innovation episode, one of the core pillars of lean development, is now more critical than ever. You simply cannot adopt AI and expect a net positive outcome without giving people the ability and the responsibility to detect problems, intervene, and continuously improve the system.

But the thing is, Morgan, when I say ability and responsibility, I don't mean giving people a dashboard and saying, "What do you mean? Of course, you're responsible." I mean fundamentally setting them up for successful ownership. That's not just tools. It's letting people decide how they produce the outcomes you expect within the [00:52:00] boundaries and norms the team has agreed on. It's planning work while recognizing that a unit of work still needs to be something a human being can review and understand, as opposed to saying, "You have forty-five minutes to build Slack."

It's supporting psychological safety so people can say, "I'm sorry, I know this looks ready, but I don't understand what this thing does or how it does it. We need more time." It's supporting the continued technical growth of employees and not devaluing knowledge of computer science or any other discipline, both in how you treat people and in how you design their work so they don't lose their skills.

Because if they lose their skills, they also lose the ability to judge the work of the AI and to act as a quality gate before release. It's understanding that prompt fatigue is real, and perhaps most importantly, it's realizing that dignity, purpose, and professional identity have repeatedly been shown throughout history to matter far more to people than money, even when money is scarce. So don't [00:53:00] assume that you can ask people to stop doing the thing they studied and worked hard to be able to do and expect them to remain engaged simply because they're getting a paycheck

Morgan VanDerLeest: it's interesting the timing of discussing this because something that, I just started doing with my teams is going back and doing a renorming exercise as if we're a new team again, as if we're a new batch of people because a lot of the underlying assumptions that we've had in the past are just different now.

And it's actually been really insightful to see what folks really want and what they want to avoid. And I think the important thing that I'd wanna communicate with folks is we as managers or leaders can't be expected to know everything and to have all the answers right now. But we as a collective group, meaning like our teams and our organizations, probably have a pretty good idea for how we can make things better and improve and work through this in a way that is good and fulfilling as people and valuable to our businesses. And it resonates with me, and I hope it does with others: Count on the people that we have, because there's a good [00:54:00] chance that with that collective intelligence, we're still able to solve things that we've wanted to solve all along

Eddie Flaisler: So true, and I love renorming. I haven't heard that word in a while

Morgan VanDerLeest: It's funny how some of the little things never go out of style.

Eddie Flaisler: Mm-hmm.

Morgan VanDerLeest: By the way, I particularly liked your Slack example. It sounds exaggerated, but it actually isn't. I've witnessed these conversations in different forums, and I always tell myself, " Let's just do the mental exercise of whiteboarding the functional requirements for this thing that you wanna build."

No non-functional requirements, no design, nothing. Just functional requirements. That alone would take a few hours. So when someone tells me they expect people to just create things within these insane timelines, it doesn't tell me that they have a high bar. It tells me that they themselves haven't put much thought into what they're trying to build.

That's not ambition, that's carelessness

Eddie Flaisler: I could not agree more

Morgan VanDerLeest: So Eddie, I think we left people with enough to think about regarding structural integrity, and we can close today's episode by discussing coherence. What does that look like?

Eddie Flaisler: [00:55:00] Well, at the risk of oversimplifying Eric's thoughtful commentary, my take is that coherence is essentially putting the right incentive system in place. One where the behaviors you want and the behaviors that get rewarded are the same. So assuming the importance of prioritizing long-term value creation over short-term transaction optimization has been internalized, you now need to make sure people are actually motivated to act in accordance with that.

So if quality matters, you promote based on that as well and not just by output volume. If collaboration matters, you don't assign adjacent teams conflicting KPIs and then wonder why they're not motivated to help each other. For example, you don't measure your API team on P95 latency while measuring your data platform team on data quality SLAs, and then you act surprised when one side wants to keep requests light while the other keeps asking for stricter schemas and stronger validation.

If you want AI to augment people rather than replace them, don't make managers feel that reducing headcount [00:56:00] is the only path to recognition. If you want sustainable systems, don't celebrate the person who saved the day while ignoring the person who prevented the emergency in the first place. I could go on

Morgan VanDerLeest: I love that

Something I want to lean back on as we've discussed the expert persona in the past, where it's this person that has a ton of useful institutional knowledge and doesn't want to share that so much with others because that's where they derive value for themselves. And especially right now, that's a thing that feels so detrimental to not just that person and their own career in the midst of all the changes going on, but just like to teams and companies as a whole. I just talked about doing this renorming exercise, and one of the things that we talked about is making sure that knowledge doesn't live within any one person, not as a managerial thing, but people can feel that, and people want to move quickly, make good decisions, not have to rely on any one person blocking the process for them, especially with AI because you can move quickly when it's possible. If you are not sharing that [00:57:00] knowledge with others, you are intentionally becoming a blocker for allowing people to move quickly. And so even more than ever before, that type of persona can be so detrimental to a team and the way that they work. So really lean into how can I make sure that the knowledge that I have is communicated well, structured well, able to be followed well.

That doesn't help just the people around you, but that helps the harness and the systems and the AI that we're building around this to operate well and to do their best along with us.

Eddie Flaisler: 100%, now more than ever

Morgan VanDerLeest: All right, Eddie I think we are done for the day

Eddie Flaisler: I think so too

Morgan VanDerLeest: Oh, so what's loop maxing?

Eddie Flaisler: Oh, that I don't even wanna talk about. And to the listeners, if you enjoyed this, don't forget to share and subscribe on your podcast player of choice. We'd love to hear your feedback. Did anything resonate with you? More importantly, did we get anything completely wrong? Let us know. Share your thoughts on today's conversation to PeopleDrivenDevelopment, that's one word, [00:58:00] PeopleDrivenDevelopment@gmail.com, or you can find us on X or Twitter at PDDPod, P-D-D Pod.

Bye everyone.

Morgan VanDerLeest: See y'all.

Creators and Guests

Eddie Flaisler
Host
Eddie Flaisler
Eddie is a classically-trained computer scientist born in Romania and raised in Israel. His experience ranges from implementing security systems to scaling up simulation infrastructure for Uber’s autonomous vehicles, and his passion lies in building strong teams and fostering a healthy engineering culture.
Morgan VanDerLeest
Host
Morgan VanDerLeest
Trying to make software engineering + leadership a better place to work. Dad. Book nerd. Pleasant human being.
AI Wars: Managerial Economics Strikes Back
Broadcast by