Free YouTube Transcribe

Video transcript

4.Usecase: Why the Best AI Models Fail in Production

Techinvest · 1,245 words · 6 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Across major technology and financial

0:02firms, the race to deploy enterprise AI

0:05is consuming massive amounts of capital

0:07and engineering bandwidth. Many

0:09technical product managers operate under

0:11a specific assumption. They believe that

0:14fielding a successful AI product is

0:16primarily a matter of selecting the most

0:18intelligent API available or training a

0:21model on proprietary corporate data. The

0:24reality in production is entirely

0:25different. Deployments utilizing the

0:28most advanced frontier models available

0:30are actively draining entire corporate

0:32budgets, failing to show measurable ROI,

0:35or losing benchmarks against generic

0:37chatbots. This exact paradox struck

0:40three highly sophisticated enterprise

0:42giants: Uber, Bloomberg, and Morgan

0:44Stanley. On the surface, their rollouts

0:47featured massive user adoption rates and

0:49models trained on billions of

0:51parameters. Beneath that surface, severe

0:54architectural vulnerabilities dictated

0:56the actual results. We can trace these

0:58vulnerabilities to three distinct

1:00constraints in production AI: agentic

1:03token economics, fine-tuning decay

1:05rates, and retrieval infrastructure.

1:08In an enterprise environment, a system's

1:10architecture dictates its survival. The

1:13raw intelligence of the underlying large

1:15language model is only a fraction of the

1:17equation.

1:18In December 2025, Uber launched Claude

1:21Code, an AI coding assistant to its

1:23engineering organization of roughly

1:255,000 developers.

1:27On paper, it was a massive product

1:29success. Adoption scaled rapidly. Within

1:324 months, 84% of engineers were using

1:34the tool, and the company reported that

1:3670% of all committed code was

1:38AI-generated. The immediate financial

1:41consequence of that success was severe.

1:43By April 2026, just 4 months into the

1:46rollout, Uber had burned through its

1:48entire AI budget for the year.

1:51Standard software budgeting operates on

1:53a linear per-seat model. A company

1:55expects to pay a flat subscription fee,

1:58which for enterprise tools typically

1:59caps at roughly 150 to 250 dollars per

2:03user each month. We are plotting the

2:05flat linear SaaS budget line against the

2:08reality of a gentic token burn, which

2:10forms a sharply exponential curve. This

2:13growth is driven by token

2:15multiplication. It starts simply. A user

2:18issues a single task prompt to an agent.

2:21The agent enters a multi-step reasoning

2:23loop. It sends a query, receives an

2:26answer, and formulates the next logical

2:28step. At every step, the agent appends

2:31the entire previous conversation history

2:34and resends that accumulating text block

2:36to the model. This recursive data

2:38transfer means a single user task

2:40multiplies token consumption

2:42exponentially compared to a simple chat

2:44exchange. Because Uber's architecture

2:47lacked built-in spending ceilings, the

2:49financial fallout scaled with the tool's

2:51adoption. Power users were suddenly

2:53consuming between 500 and 2,000 dollars

2:56per month in tokens. To compound the

2:58issue, Uber's executives admitted they

3:01could not draw a clear line between this

3:03massive internal token spend and any

3:06corresponding improvement in

3:07consumer-facing business results. High

3:10adoption metrics are irrelevant if the

3:12underlying unit economics are flawed.

3:14Failing to architect for non-linear

3:16token economics turns a successful

3:18product rollout into an uncontrolled

3:20financial liability.

3:22Two years earlier, in March 2023,

3:25Bloomberg announced a different

3:26approach. They decided to build a custom

3:29language model trained specifically for

3:31finance. The investment required massive

3:33upfront capital and compute. Bloomberg

3:36GPT featured 50 billion parameters

3:39trained on a mix of public data and over

3:41360 billion proprietary financial

3:44tokens. The training process alone cost

3:46an estimated 3 to 10 million dollars.

3:49The strategic premise was common in

3:51enterprise AI at the time. Pairing

3:53proprietary corporate data with

3:56dedicated from-scratch training

3:58inherently creates a durable competitive

4:00advantage. Upon release, the premise

4:02seemed validated. Bloomberg GPT

4:05successfully outperformed the

4:07similarly-sized open-source models

4:09available at that time on

4:10finance-specific tasks.

4:12However, Bloomberg's original paper

4:14omitted a critical data point. They did

4:17not directly benchmark their custom

4:18model against frontier models like

4:20GPT-4.

4:22This chart displays findings from

4:24independent researchers at Queen's

4:25University who ran the direct comparison

4:28on the FinQA benchmark. Bloomberg GPT

4:31scored 43% while GPT-4 reached 68.79%.

4:36GPT-4 achieved a nearly 26-point lead on

4:40a specific financial benchmark,

4:42possessing zero specialized financial

4:44training. This performance gap exposes

4:47the flaw in custom training, a concept

4:49known as fine-tuning decay.

4:52Custom training freezes a model's

4:53knowledge as a static snapshot in time.

4:56This timeline graph illustrates the

4:58decay.

4:59The static state of Bloomberg's custom

5:01model capability forms a flat horizontal

5:04line. Against that, we plot the

5:06continuous advancement of

5:07general-purpose models as an

5:09upward-trending curve.

5:11As general models advance, the

5:13capability gap that originally justified

5:15a multi-million-dollar custom training

5:17investment rapidly narrows and

5:19eventually vanishes entirely.

5:21Proprietary data is valuable, but baking

5:24it directly into a model's weights is

5:26brittle.

5:27Over-investing in static custom model

5:29training is a losing bet against the

5:31relentless evolutionary pace of

5:33general-purpose APIs.

5:35Morgan Stanley provides the counter

5:37model. They successfully deployed an

5:39AI-powered internal research assistant

5:42for their wealth advisors. Rather than

5:44spending millions training a model from

5:46scratch like Bloomberg. Morgan Stanley

5:48directed their engineering resources

5:50toward meticulously organizing their

5:52internal data. The scale of this

5:54knowledge base was immense. They

5:56launched with over 100,000 internal

5:59research documents, eventually scaling

6:01the corpus past 350,000 files. Asking a

6:05generic LLM to rely on its internal

6:07memory to recall specific financial

6:09details guarantees high hallucination

6:12rates. The model will confidently invent

6:14facts it cannot retrieve. This schematic

6:17maps out their solution, a retrieval

6:19augmented generation or rag pipeline. In

6:22the embedding phase, a user query is

6:24converted into a mathematical vector to

6:26scan the massive document database. Next

6:28is the critical retrieval action. The

6:31system isolates only the highly relevant

6:33internal documents based on strict

6:35mathematical vector matching, ignoring

6:37the rest of the corpus.

6:39Finally, in the augmentation phase,

6:41these retrieved factual documents are

6:43packaged alongside the user's original

6:45prompt before the query ever touches the

6:48large language model.

6:49Morgan Stanley spent months curating

6:51this pipeline. They used human experts

6:54to rigorously test responses, explicitly

6:56evaluating factuality and hallucination

6:59rates long before scaling the tool to

7:00users. This data readout highlights the

7:03exact metric that proves the

7:04architecture's value. Retrieval

7:06efficiency reportedly rose from roughly

7:0820% to 80% after the rag system

7:11launched. That massive jump in system

7:13accuracy was achieved in entirely in the

7:15vector database and retrieval layer,

7:17independent of the baseline intelligence

7:19of the LLM. By establishing this robust

7:22predictable pipeline, Morgan Stanley

7:25safely expanded their architecture in

7:262024, adding a second automated tool to

7:29generate meeting notes and follow-up

7:31actions from client conversations.

7:34In production environments, the

7:35underlying large language model is

7:37interchangeable. The data retrieval

7:39infrastructure is the most reliable

7:41lever for system accuracy and measurable

7:44ROI.

7:45Treating enterprise AI deployment like

7:47standard software procurement guarantees

7:49failure at scale. The budget models and

7:52infrastructure requirements are

7:53fundamentally different. Uber's rollout

7:56yields the first rule. You must engineer

7:58strict token metering and hard spending

8:01caps to control the non-linear cost of

8:03agentic reasoning loops.

8:05Bloomberg's multi-million dollar

8:07training yields the second rule. You

8:09must assume the baseline intelligence of

8:11general APIs will quickly outpace static

8:14fine-tuned custom models. This unified

8:17system architecture diagram outlines the

8:20mandate for technical product managers.

8:22You need a token metering firewall to

8:24limit cost, a dedicated rag and vector

8:27database retrieval layer to secure

8:29facts, and a generic LLM API block to

8:32process text. Morgan Stanley success

8:35confirms the final rule. Direct the

8:37majority of engineering resources and

8:39evaluation frameworks toward that

8:41retrieval layer, rather than custom

8:43model training.

8:44True competitive advantage in enterprise

8:46AI does not lie in the specific

8:48intelligence of the model you choose

8:50today. It lies in the resilient, highly

8:52controlled pipeline you build to manage

8:54it tomorrow.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.