Full transcript
0:00In the absence of AI and robotics, we're
0:02actually totally screwed.
0:03>> We are working to build tools that one
0:05day can help us make new discoveries and
0:06address some of humanity's biggest
0:08challenges, like climate change and
0:10curing [music] cancer.
0:11>> Hi, welcome to another episode of
0:12ColdFusion.
0:14Here's a question. [music] How can AI be
0:16disrupting the job market, but also be
0:18losing billions of dollars at the same
0:19time?
0:20Well, this video will answer that.
0:22>> [music]
0:22>> The truth is, while AI helps make some
0:25jobs easier, when compared to a human,
0:27it performs worse a whopping 96.25%
0:31of the time, which basically means,
0:33given AI 10 tasks, and it will perform
0:35at least nine of them worse than when
0:37compared to a human.
0:38That's at least according to a new
0:40study.
0:41It's such an interesting finding and
0:42begs the question, why has no one
0:44systematically compared how well AI does
0:47versus a human who's done exactly the
0:49same job? All previous benchmarks have
0:51been simulated human work, not real
0:53generalized work.
0:55The results from the team of researchers
0:57who did the study makes one think, maybe
0:59the true value of consumer AI isn't
1:01hundreds of billions of dollars, but
1:03orders of magnitudes less. I'm not
1:05saying that all AI sucks. This study is
1:08just a general reminder that AI is a
1:10time-saving tool and not a replacement.
1:13Just maybe, the economy is valuing it
1:16too highly when it comes to near-term
1:17capabilities.
1:21In this episode, we'll take a look at
1:23the study in detail and discuss [music]
1:25what it all means.
1:28You are watching ColdFusion [music]
1:30TV.
1:32So, the synopsis of the study was
1:34straightforward enough. Give paid jobs,
1:36already completed by real people, to AI
1:38models, and then see how well the
1:40results compare. Once the AI completes
1:42the tasks, humans evaluate the results.
1:45The researchers called this method the
1:47Remote Labor Index or RLI.
1:49>> [music]
1:49>> It's so simple. Most of us use a
1:51computer to do modern work, right? So,
1:53why not just directly compare how well
1:55AIs compete on a professional
1:56computer-based job.
1:58The jobs to be completed were real ones
2:00from the freelancer site Upwork, a site
2:02where you pay remote workers to complete
2:04any given task.
2:06The jobs were varied from video
2:07creation, computer-aided design,
2:09>> [music]
2:09>> graphic design, game development, audio
2:12work, architecture, and more.
2:15Both humans and AI were given the same
2:17brief and any attached files that were
2:19necessary for the job. For example, an
2:21Excel spreadsheet of data or
2:23instructional images.
2:26The AI models were tested on 240 jobs,
2:28each paying $630 on average.
2:31So, how did they perform?
2:34The performance was abysmal. The best AI
2:36was Claude Opus 4.5 with a 3.75% success
2:40rate when it came to producing work of
2:41an acceptable quality. You heard that
2:43right, a 96.25%
2:45failure rate was the best performer.
2:48Interestingly, Gemini was the loser with
2:50a 1.25% success rate. Now, Claude Opus
2:544.6 might score 5% better, but that's
2:56still a 91% failure rate. When these
2:58scores get to 35% or 40%, then we can
3:01talk.
3:02So, a couple of things to note. The
3:04original paper used AI models that were
3:056 months or so old, but their website
3:08has up-to-date results, which are the
3:09scores that I'm referring to in this
3:11episode. I'll leave a link for the
3:12website below.
3:14So, where exactly did the AI systems
3:16fail?
3:17Well, first we need to define exactly
3:18what failure means. Failure counts as
3:21not performing a task at or better than
3:23a human level. This is specifically in
3:25the context of a freelancing
3:26environment, an environment where people
3:29actually pay money directly for the
3:30work.
3:31With that in mind, the paper lists four
3:33main failure points for AI systems.
3:36Number one, sometimes the AI would
3:38produce, quote, [music] "corrupt or
3:40empty files" or deliver work in
3:42incorrect or unusable formats.
3:45Number two, [music] AI, quote,
3:47"frequently submitted incomplete work
3:49characterized by missing components,
3:51truncated videos, or absent source
3:53assets. For example, a video of 8
3:56seconds when an 8-minute video was
3:57required.
3:59Number three, another one was quality
4:01issues. Quote, "Even when agents produce
4:03a complete deliverable, the quality of
4:05work is frequently poor and does not
4:07meet professional standards." End
4:09[music] quote.
4:11And finally, number four,
4:12inconsistencies with AI-generated work.
4:16This includes a house's appearance
4:17changing across different 3D views or
4:19digital floor plans that don't match the
4:21supplied sketches. It's all very
4:23interesting. So, for years now, we've
4:25been told that AI is going to replace
4:27humans everywhere, but the truth is, we
4:30are nowhere near that point, at least
4:32not yet, anyway.
4:33>> [music]
4:34>> So then, where did the AI succeed?
4:37Success would mean that the AI does the
4:38same work at the same quality or better
4:41quality than human output. They note
4:43that AI was proficient in creative
4:45ideas, like audio and image-related
4:47work, along with writing, data
4:49retrieval, or web scraping. And that
4:51kind of checks out. The success of
4:53OpenClaw attests to the latter, too. And
4:56AI images and audio are already good
4:58enough to fool a lot of people.
5:00Advertisement and logo creation was
5:02another successful area. It's also no
5:04surprise that AI was good at report
5:06writing and generating simple code for
5:08an interactive data visualization.
5:11Competent video generation is coming
5:13very shortly. Just take a look at
5:15SeeDance 2.0.
5:20>> [music]
5:27[music]
5:33>> You didn't know.
5:35>> [music]
5:39>> I didn't know.
5:48>> [music]
5:58[music]
6:05>> So, the main takeaway is AI is pretty
6:07good at some things, but horrendous for
6:09general work.
6:11But, what else do we learn?
6:13This paper exposes a lot, much of it
6:15negative, but it does show that the RLI
6:18format is a very useful measure of AI
6:20performance in the real world.
6:22Reason being, current-day benchmarks
6:24aren't reflective of real-world
6:25performance.
6:27As the paper puts it,
6:28>> [music]
6:28>> quote, "While AI systems have saturated
6:30many existing benchmarks, we find that
6:33the state-of-the-art AI agents perform
6:35near the floor on RLI." End quote. I
6:37found the study to be very robust, by
6:39the way. So, I'll leave a link to it
6:40below.
6:41>> [music]
6:42>> According to this study, AI may impact
6:44jobs with lots of language requirements,
6:46audio, simple advertising, or data
6:49retrieval, but human oversight is still
6:51needed. A PwC report found that the
6:54majority of CEOs see no financial
6:56returns from AI. Upper management and
6:58CEOs just command workers to use AI and
7:01expect it to all work. For AI to work
7:03within a corporation, there needs to be
7:05a planned and skilled implementation of
7:07the technology with the knowledge of its
7:09shortcomings, and that doesn't happen a
7:11lot of the time. Gartner predicts that
7:12by next year, half of the companies that
7:15fired workers for AI are going to hire
7:16them [music] back.
7:18Also, 9 months ago, Microsoft proudly
7:20proclaimed that 30% of their code was
7:23written by AI, and since then, we have
7:25seen some of the worst software issues
7:26at the company in its history.
7:28Now, it's obvious that AI is disruptive,
7:31and some jobs will be lost to the
7:32technology. For example, diffusion
7:34models are proficient in the visual
7:36arts, as you saw earlier, but as for
7:38LLMs in the general workforce, this
7:40study indicates that job losses could be
7:42a lot less.
7:43>> [music]
7:43>> The AI space does move fast, so I could
7:45be wrong, but that's how things are
7:47looking today in early 2026. To sum up
7:49the job prognosis in one line, if you're
7:51a software engineer, set up a business
7:53that fixes vibe-coded apps, and you'll
7:55make a lot of money. I think the thing
7:56[music] is, artificial intelligence
7:58really is going to transform the world,
8:00like, in ways we can't even imagine, but
8:03it's not going to do it now, not [music]
8:04with this technology. My favorite
8:06example of this is one trains them on
8:08the whole internet, so they get access
8:10to a lot of written rules of chess and
8:12lots of games of chess, and they still
8:14make illegal moves. They never really
8:16[music] abstract the model of how chess
8:19works. That's just so damning. You would
8:22not be able to learn chess after seeing
8:24a million games, reading the rules on
8:26Wikipedia and chess.com. Just making it
8:29bigger is not going to solve these
8:30problems. We need to do foundational
8:31research. That's what I was saying for
8:33the last 5 years. What is intelligence
8:35of the problem is is to understand your
8:37world, and um
8:40>> [music]
8:40>> Reinforcement learning is about
8:41understanding what your world where is
8:43large language models are about
8:45mimicking people. Doing what people say
8:47you should do. They're not about
8:49figuring out what to do. Just to mimic
8:51the the what people say is not really to
8:53build a model of the world at all, I
8:54don't think. So, I'm not saying that AI
8:57will never work or it's not generally
8:59useful already. There will be some
9:01narrow AI products that work really
9:02well. I'm just warning that there's a
9:04significant financial risk in the
9:06current AI space. The investment ethos
9:08and the rollout of AI everywhere might
9:10be misallocating hundreds of billions of
9:12dollars.
9:14Even in the medical field, Reuters just
9:16reported that the FDA has received 100
9:18reports of AI malfunctions, botched
9:20surgeries, and misidentified body parts.
9:23In a few cases, a lawsuit alleges that
9:25the AI misinformed the surgeons on the
9:27locations of their instruments, causing
9:29one to mistakenly puncture the base of a
9:30patient's skull, and causing strokes
9:33from the damage to a major artery in two
9:34others. We don't need to put AI in every
9:36field. It's just not ready yet. Again,
9:39in some fields like coding, high maths,
9:41and writing, AI is pretty good and can
9:43make jobs a lot easier, but we can't
9:45pretend like it's going to replace
9:47everyone perfectly right now. Now, I was
9:49going to stop the video here, but just a
9:51couple of personal thoughts. Back in
9:522016 when I started covering AI, it was
9:55fun and fascinating to see how these
9:57things worked, but ever since the big
9:59money started coming in, the hype has
10:01just gone off the charts.
10:03>> [music]
10:03>> CNBC just reported that companies like
10:05Anthropic, Google, and Microsoft have
10:08paid individual content creators
10:10$400,000 to half a million dollars each
10:13to promote their AI models.
10:15Now, brand deals are fine, but if the
10:17current generation of AI was as
10:19revolutionary as being advertised,
10:21[music] they wouldn't need to spend so
10:23much money to convince us. It's a
10:24jarring disconnect. One last thing.
10:27We're fooled into thinking those
10:28machines are intelligent because they
10:30can manipulate language, and we're used
10:31to the fact that
10:33people who can manipulate language very
10:35well are implicitly smart, but
10:38we're being fooled.
10:39Um now, they they're useful, there's no
10:42question. They're great tools, like, you
10:44know, computers uh have been for the
10:46last five decades five [music] decades.
10:48But, let me make an interesting
10:50historical point, and this is maybe due
10:51to my age. Uh
10:54there's been generation after generation
10:57of AI scientists
10:59since the 1950s claiming that the
11:03technique that they just discovered
11:05was going to be the ticket [music] for
11:07human-level intelligence. You You see
11:09declarations of Marvin Minsky,
11:12Newell and Simon, um you know, uh
11:15Frank Rosenblatt who invented the
11:17[music] perceptron, the first learning
11:18machine in 1950, saying like, within 10
11:20years we'll have machines that are as
11:22smart as humans. They were all wrong.
11:25This generation with an LLM is also
11:27wrong. I've seen three of those
11:29generations in my lifetime, okay?
11:32Um,
11:33so, you know, it's it it's just another
11:36example of being fooled. That's Yann
11:38LeCun, the creator of convolutional
11:40neural networks. He's been outspoken in
11:43saying that the current AI architecture
11:44is reaching its peak. He thinks that
11:46throwing more data and power at the
11:47problem isn't going to solve it. And I
11:50think that's what the early data is
11:51showing us. It's called the scaling
11:53problem, and it's a large part of my
11:55upcoming video about how OpenAI is in
11:57big [music] trouble. When it's complete,
11:59I'll leave a link to that episode below.
12:00So, be sure to check it out after this.
12:03Anyway, [music] that's about it for me.
12:05You've been watching ColdFusion. Let me
12:06know your thoughts. I'm sure the comment
12:08section will be very very full of very
12:10good [music] discussions.
12:12Anyway, that's it. My name's Dagogo, and
12:14I'll see you again soon for the next
12:16episode. Cheers, guys.
12:17>> [music]
12:17>> Have a good one.
12:28>> [music]
12:36[music]
12:42[music]
12:45>> ColdFusion.
12:46It's new thinking.