Full transcript
Intro
Title: Navigating AI Project Hurdles: Insights and Expertise from ReflexAI
0:07>> Our industry research shows that 80%
0:10or more of AI projects fail
0:12to reach meaningful production deployment
0:15and development delivering measurable business value.
0:18This is highlighting the operational
0:19and organizational challenges in scaling beyond prototype.
0:22My name is Paul Nashawaty, I'm the practice lead
0:24and principal analyst for the AppDev practice,
0:27and I'm joined by John.
0:28John, how are you doing today? You're from
0:29ReflexAI, correct?
0:31>> Doing well. Thank you. Thank you for having me.
0:34>> Yeah. John, tell me a little bit about yourself
0:36and about ReflexAI.
0:37>> Yeah, well, thanks for having me on. I am John Callery.
0:42I am the co-founder and chief product
0:45and technology officer at ReflexAI,
0:48and we build AI-powered simulations
0:50and quality assurance software
0:52that improves human performance in complex conversations.
0:57The company grew out of work my co-founder, Sam Dorison,
1:01and I did when we were both leaders at The Trevor Project,
1:04and there we built AI-powered conversation simulations
1:08to train crisis counselors at scale,
1:11and that work ended up getting major recognition,
1:14including Times 100 Best Inventions
1:17of 2021.
1:19One through line from
1:21that experience is AI can accelerate experimentation
1:25dramatically, but the systems you rely on in the real world
1:28still require strong engineering discipline
1:32and operational rigor,
1:33and I think that's really the lens that I bring
1:36to this conversation about vibe coding
1:38and scaling AI in enterprises.
1:41>> Well, John, it's super relevant.
1:42I mean, it makes a lot of sense.
1:44We see a lot of organizations moving from vibe coding
1:47and AI experimentation really to production
1:49to scale deployment.
1:51But this is a challenge that
1:53no longer can just model access,
1:54it really requires operational discipline.
1:56Because when we're looking at things like governance,
2:00compliance, regulations,
2:01and control, enterprises really must have an established new
2:05engineering standard that manage AI-native tech debt,
2:08ensure governance, but also prevent uncontrolled complexity
2:12really from compounding across code, data,
2:15and automation layers.
2:16This is really kind of an extension of what we were seeing
2:20with shadow IT and shadow development,
2:24but this is now becoming shadow
2:26code and application development.
2:28At the same time,
2:29we're seeing innovation can no longer be confined
2:32to these technical teams
2:34because they just don't have the amount of
2:37resources in order to do it for professional developers.
AI Innovation and Governance: Bridging Creativity and Responsibility
2:40So we have to let citizen developers in product
2:43and operations and finance
2:44and go-to-market functions really help
2:46build products of their own.
2:47So, John, what are your thoughts on that?
2:48I mean, it really is kind of an area
2:51that's expanding from the professional
2:52developer to the citizen developer.
2:55>> I think democratizing AI is essential
3:00because a lot
3:01of the best AI opportunities come from non-technical teams.
3:05If you think product, design, operations,
3:08and go-to-market, those are folks who are often closest
3:11to real workflow problems.
3:14The risk is exactly what you said.
3:16If people are experimenting independently,
3:20you can get kind of fragmentation
3:22and inconsistent standards in shadow AI,
3:25but the solution isn't to stop experimentation.
3:29The solution is to make experimentation safe and structured.
3:34One of the reasons I feel really strongly about this is
3:37we're living it right now.
3:38This week, we intentionally changed how we work
3:42so every role in the company builds real AI fluency as part
3:46of their daily workflow, not just as a side hobby,
3:50and the point of this is to remove bottlenecks.
3:52If only engineers can use AI effectively,
3:56then product context, customer nuance,
3:59and operational insight don't make it into the system,
4:02you end up with output that's fast
4:05but lacks user-centered soul
4:07that it actually needs to succeed.
4:09Practically speaking, the organizations
4:13that do this well provide shared platforms, clear policies
4:17and support, and they also create a culture where
4:21people can share what worked and what didn't without fear.
4:25I just want to be clear, this is not about replacing people,
4:27it's about giving people leverage
4:30and raising the bar of quality across the organization.
4:35>> Yeah, absolutely.
4:37I think that there's this misnomer of AI is going
4:41to come in and take a new job.
4:43I don't think that's really the case.
4:45I mean, it's an operational tool.
4:47Because you can do math on pen and paper,
4:50but you can use a calculator.
4:51It's kind of a similar approach here.
4:53But let's talk about something you talked about,
4:54which is bridging the vibe coding to production reality.
4:58The gap isn't about model quality,
5:00it's about operational rigor.
5:02When we look at security, observability, versioning,
5:04governance, it really comes down
5:06to prototype success often masked with missing fundamentals.
5:11You think things like data quality or latency constraints
5:14or cost controls or compliance.
5:16Those are things that organizations must
5:18understand and know.
5:20And then when you have AI system
5:21that have probabilistic kind of production environments, it
5:27really demands guardrails, right?
Title: Navigating AI: Strategies for Safe and Inclusive Innovation
5:29Because if you don't have those guardrails in place,
5:31then it's the kind of a little bit of chaos.
5:33And I get it that when...
5:36You kind of want to use the words of former Andy Grove from
5:41Intel who said like, "Let the chaos...
5:43and then rein it in." I guess that's true in the context
5:46of moving production requires evaluation frameworks,
5:51it requires clear ownership,
5:52it also requires production grade data pipelines.
5:55So, John, when I'm looking at this, we're seeing a lot
5:57of vibe coding and AI experimentation,
6:00but what fundamentally breaks when organizations try
6:04to move these prototypes into production?
6:07>> Vibe coding is incredibly powerful for exploration.
6:11You can iterate with an LLM quickly,
6:15you can test ideas,
6:16and get to early signal way faster
6:20than traditional development.
6:21The problem is that prototypes hide operational complexity.
6:27A demo can look great while skipping all of the fundamentals
6:31that only surface later.
6:32So once you go to production,
6:35the requirements completely change.
6:37Now, you need reliable data flows,
6:41like you said, automated testing, observability,
6:44and security controls,
6:46and really clear ownership over
6:48how the system actually behaves.
6:51Like you mentioned, another piece is
6:53that the AI systems are probabilistic.
6:55Traditional software is mostly deterministic,
6:58but AI behavior can shift based on everything from context
7:03to model changes to just subtle differences
7:05and input, so production environments really need those
7:08guardrails and evaluation.
7:10It's more than just a working demo.
7:13I think even here at ReflexAI,
7:15we've seen this firsthand in high stakes environments.
7:18When the work also affects real human outcomes,
7:22you can't just rely on something that merely seems to work.
7:24You really need evaluation frameworks and monitoring
7:28and disciplined release practices.
7:31There's also a hype dynamic right now
7:33where people are talking about shipping features at
7:36incredible speed.
7:38Even the frontier AI labs talk about the
7:43tools that they've built and shipped in weeks.
7:44But if you've
7:49actually used a lot of these AI apps
7:51or looked at their status pages over time, it's a reminder
7:55that operating AI systems reliably is still very hard.
7:59So if the companies
8:00who are building the models themselves are constantly
8:04running into stability
8:06and performance of even their own apps, it
8:11just shows you how big the gap really is
8:13between rapid prototyping and dependable production systems.
8:17So the takeaway isn't like, "Don't move fast," it's,
8:21"Systematize velocity,
8:22and move fast with those guardrails so speed compounds,
8:27instead of creating a lot of fragility in the system."
Concluding Thoughts and Next Steps
8:30>> Yeah, there's no doubt.
8:31I think that the other thing to kind of think about here,
8:34John, is to manage the AI-native technical debt.
8:37You mentioned shadow AI. I have said shadow IT.
8:40But I think when we look at prompt sprawl
8:42and agent sprawl, that's really the new kind of sprawl
8:46that we're seeing in the shadow world of AI or IT.
8:50But model upgrades can silently break workflows,
8:52like you were talking about.
8:55It hides the complexity,
8:57and versioning discipline is now critical, right?
9:00AI really can introduce data debt, context debt,
9:04alignment debt.
9:07Emerging best practices are really kind
9:09of implementing things like AI governance layers
9:11and model registries
9:13and prompt lifecycle management as well
9:15as AI observability stacks.
9:16So the question really isn't, "Can we build it?
9:19" It's, "Can we sustain it at the end of the day?
9:22" When I look at that, AI-
9:24native technical debt is becoming a real issue.
9:27It really is because there's a lot of development out there.
9:30But how is it different from traditional software debt,
9:32and what is new about engineering standards that help...
9:35that barely emerge to manage this debt
9:38that's coming a whole lot?
9:42>> Traditional technical debt usually comes from messy code,
9:47shortcuts in architecture, or missing tests,
9:49and it's painful,
9:51but it's familiar territory for the industry.
9:54AI creates new categories of debt
9:58that are a lot less visible.
9:59Like you mentioned, you get prompt sprawl, agent sprawl,
10:03inconsistent context handling,
10:05and systems that quietly break when the model
10:09of prompt changes, and so
10:12at the core difference here is the failure mode.
10:15Because traditional software tends
10:17to fail deterministically, you can reproduce the bug.
10:20AI systems are probabilistic, so
10:23that same workflow can behave differently depending on the
10:28context, on data drift,
10:30or how the model's behavior shifts over time.
10:34What this means is that you really need
10:36to bring new standards.
10:38Versioning can't just apply to code anymore.
10:40It has to also cover prompts, the models, evaluation suites,
10:45and the data context that drives behavior,
10:49and you also need continuous evaluation.
10:51So the question isn't, "Did it work once?
10:54" It's, "Is it still working across scenarios at
10:58scale and under real conditions?
11:01" And I think the organizations
11:03that handle this well treat AI features like real
11:06products with lifecycle management.
11:07It's not just one-off experiments
11:10that are stitched together into a fragile foundation.
11:13>> Yeah, that makes a lot of sense.
11:16I also want to kind of go back to something
11:17that you said earlier, democratizing, right?
11:20You talk about democratizing the
11:22AI deployments by creating and such.
11:24I want to call it democratizing innovation,
11:26but really democratizing innovation without losing control.
11:29And looking at it in the context of non-technical teams,
11:33they're really closest to their own business problems.
11:35That's where AI opportunity often lives
11:38because they're trying to solve a problem.
11:40But the risk is decentralizing experimentation without
11:43guardrails causes...
11:44obviously, they can go off the rails real quick.
11:48So what we see in best practices,
11:49we see winning organizations really are creating safe
11:52sandboxes, internal AI platforms, curated model access,
11:56and clear usage policies.
11:57Those are the ways to put those guardrails in place.
11:59But AI literacy is becoming a competitive advantage,
12:02not just a developer skill,
12:04and the shift from citizen developer
12:06to citizen AI operator is really a model that's happening.
12:09I'm seeing that.
12:11In a lot of organizations, I'm sort of looking at it.
12:14So, John, when we look at democratizing AI innovation,
12:18it sounds powerful,
12:19but how do you really bring non-
12:21technical teams into the fold without creating
12:24chaos or shadow AI?
12:27>> Yeah. Like I mentioned, we're kind of going
12:30through this internally,
12:31and what we did was we started with practical
12:36hands-on training that is focused on prompt fundamentals
12:42and workflow patterns
12:43and then immediately moved to applied AI work
12:47where teams
12:49beyond engineering are producing real
12:52deliverables on real tasks.
12:56We also set principles to keep that safe.
13:00Those principles include individuals on the output, review
13:04of the output is non-negotiable, we share learnings openly
13:09in a shared Slack channel or in team meetings,
13:12and then iterate on
13:14that over time rather than waiting for perfection.
13:18And so that's how we've tackled that internally, especially
13:21as we've wanted to upskill the entire organization on
13:25AI and make it a useful tool in everybody's workflows
13:29beyond engineering.
13:31And I think, if we ask ourselves like,
13:35"Why do engineers seem ahead of everyone else?
13:37" I think engineers are starting ahead
13:40because code is structured
13:42and the feedback loop is immediate.
13:44You write something, you run it,
13:46and then you see what happens,
13:47but the next wave is cross-functional.
13:50The biggest unlock happens when domain experts can express
13:55workflows clearly and consistently
13:56and AI becomes
13:58that daily overlap across every role in the organization.
14:02>> Yeah. No, I really like where this conversation's going
14:04because it really kind of drives the point home of culture,
14:07bottoms-up versus a top-
14:09down AI strategy when you're looking at culture.
14:11>> Yeah. - So bottoms-up really drives
14:14that innovation speed, like you were talking about.
14:16Tops-down really ensures risk management and alignment.
14:20But when you look at it, too much of a bottom-
14:22up approach can potentially equal fragmentation,
14:25and too much of a top-down approach is bureaucracy-
14:29installed adoption, right?
14:30We kind of call it what it is.
14:31But I think the emerging model into where you're going
14:35with this is centralized AI platform strategy
14:39and decentralized use case
14:42innovation is really key.
14:44But I also think,
14:45and you kind of nailed this in my opinion,
14:48executive sponsorship is critical,
14:50but so is grassroots momentum, so you need to kind of have
14:53that balance between the two.
14:55And when we look at it, there's often tension between
15:01the bottom-up experimentation and the top-down mandates.
15:04So, John, what's the right balance
15:06to scale AI responsibility really while
15:09maintaining that velocity?
15:10And it's kind of on both sides, right?
15:13>> Yeah. I mean, this tension shows up everywhere right now.
15:17Bottoms-up experimentation creates speed and creativity,
15:21and that's where the surprising use cases are coming from.
15:25But too much bottom-up leads to fragmentation, so
15:30different tools, different standards, duplicated work,
15:35and risk exposure you potentially don't even know about.
15:40On the flip side, top-down leadership is necessary
15:43for alignment and risk management,
15:46but mandates don't really work either.
15:50A target like increased usage of AI
15:55by 20% is a metric, it's not a strategy.
15:59And so the best pattern we've found is a hybrid model,
16:03centralize the infrastructure and the guardrails,
16:05but decentralize use case discovery.
16:09That's also what we're doing internally as well.
16:12Leadership sets the direction, the principles
16:15and the safety boundaries,
16:16and then teams experiment out loud within those constraints
16:20to share what works.
16:23I think when you get that balance rate,
16:24you don't kill velocity.
16:26It's, again, kind of going for that velocity compounded.
16:31When I think about what should executives actually do,
16:34I think it's fund
16:36and prioritize shared infrastructure, define governance,
16:41and align the teams around the right outcomes,
16:45but don't expect the best use cases to emerge...
16:50actually, rather, expect the best use cases
16:53to emerge from the people who are closest
16:55to the work and not from the top.
16:58I think you could also avoid stifling innovation
17:00by making the rules about safety and quality
17:04and not about controlling ideas,
17:06and so you want people to try things freely inside
17:10of a sandbox, as you mentioned, where failure is really safe
17:13and also those learnings get shared.
17:16>> Yeah, for sure. But it also comes down
17:18to organizational readiness and orchestration.
17:20I mean, when we look at this, the high success
17:22of AI successes, it increasingly is about cross-
17:25functional orchestration, right?
17:26Whether you have engineering, security, legal, product,
17:28and finance, they're all working together,
17:31and that's a big factor.
17:32But if you have an AI change in cost model,
17:36which has variable inference costs versus fixed
17:38infrastructure, you have procurement
17:40and government processes that they have to evolve,
17:43and then there's also measurement for success
17:45because it's from that deployment that we're talking about,
17:49and this is something that is really near
17:51and dear to my heart when it comes to like the SDLC
17:54and across delivery, you have
17:56to look at things like productivity gains
17:58and cycle time reduction
18:00and revenue influence and risk reduction.
18:03Those are the areas that I think are really going
18:05to help measure the balance, to your point there.
18:11But I guess I would ask this question here, it's like,
18:14today, if you were advising a board
18:18of leading indicators that really would help drive
18:22or tell whether the organization is truly optimized
18:25or operationalizing AI, what would you say is
18:31the optimized
18:32and operationalized AI versus just experimenting with it?
18:35What would you say that to the board
18:38from those measurement perspective?
18:40Because I think that's important for the
18:41audience to understand.
18:43>> Yeah. I think
18:44that the biggest trap is mistaking
18:47pilots for progress.
18:49It's really easy to create impressive demos,
18:53and it doesn't mean that the organization is
18:55generating real value.
18:56So I think the clearest leading indicator
19:02is whether AI is actually embedded into real workflows,
19:07and are teams using it to produce real outputs every week,
19:12and are those outputs improving speed, quality
19:16or decision making.
19:19Another indicator is measurement.
19:22Mature organizations track things, like you mentioned,
19:25productivity gains, cycle time reduction,
19:30risk reduction, and it's not just the number of AI projects.
19:35And I think a third is governance maturity.
19:38Are there evaluation pipelines? Is there monitoring?
19:41Is there version discipline?
19:44Is there clear ownership in the organization?
19:46And then I think there are real business outcomes.
19:49In the training context, for example, like our customers,
19:53we've seen things like faster onboarding, higher confidence,
19:58and measurable improvements in communication skills when the
20:01AI is genuinely embedded in how teams learn and improve.
20:05I think boards should be asking themselves three questions,
20:08"Where is AI creating measurable value?
20:12How are we managing the risk?
20:14And do we have AI-native engineering standards
20:18that make it sustainable?
20:19" And so if I had to pick, I think, one leading indicator,
20:24it's workflow penetration.
20:25If AI is showing up naturally in the cadence of work
20:29and people are sharing repeatable patterns, it's real.
20:32If it's only demos and isolated pilots, it's not.
20:35>> Yeah. Yeah, no, that makes a lot of sense.
20:37I think there's a lot that we talked about
20:39during this presentation, a lot
20:41for the audience to think about.
20:43It's new to a lot of organizations,
20:45but it's also something that they need to be aware of.
20:47It's a competitive advantage,
20:49and those organizations that don't take advantage are going
20:51to be kind of left behind, honestly.
20:53But, John, thanks for being on.
20:55What would you leave with the audience on
20:57where they can learn more
20:58about what we're talking about today?
21:00>> One thing I just want to add is one thing
21:03that we've learned across industries is that
21:06as automation takes on routine work, the conversations
21:11and decisions that reach humans get more complex,
21:15more nuanced, and even more higher stakes.
21:20So the opportunity isn't only, "Where can AI replace people?
21:25" The real opportunity is,
21:27"How do we make humans dramatically better at the
21:31moments that matter the most?
21:32" And to do that responsibly, you need both sides.
21:35You need rapid experimentation
21:37and strong operational discipline
21:40so the systems you build are actually trustworthy at scale.
21:45I think one other item is if you'd like
21:48to learn more about ReflexAI, I really encourage you to go
21:51to reflexai.
21:52com and feel free to reach out. We'd love to chat with you.
21:56>> John, thanks for being on today.
21:57It's really an interesting conversation, really a lot
22:00for the audience to think about regardless
22:02where they are on their journey.
22:03It's definitely something for them to consider.
22:05And a big thank you to all of you who've tuned in.
22:07We really do appreciate you being part
22:09of the AppDevANGLE community.
22:10But for now, that wraps up this episode,
22:12but we'll be back next Wednesday
22:13with another conversation diving into the tools, trends,
22:15and talent shaping the future of application development.
22:18So when you're deploying at the edge, building with AI,
22:21or modernizing your cloud stack, we've got you covered.
22:23Be sure to follow us on social if you have any thoughts,
22:25questions, or just want to connect.
22:27Until next time, stay curious and stay building.