Navigate Select ESC Close

GPT 5.2 is the first HUMAN LABOR replacement

2025-12-12 Education
29.4k
1.1k
412
Wes Roth
Wes Roth
323.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Launch your site for free at https://framer.link/WesRoth Use code WESROTH for a free month on Framer Pro. Build a site that looks hand-coded. Without hiring a developer. ______________________________________________ VIDEO SUMMARY In this video, we test the newly released GPT 5.2 Pro and its "extended thinking" capabilities. We push the model to create complex 3D simulations—including a spherical Conway's Game of Life and a destructible city game—in a single prompt. The results show a model that acts less like a chatbot and more like a remote engineer, taking up to an hour to reason through code architecture before delivering a final project. We also break down the new "GDPval" benchmark. unlike traditional tests, this evaluates AI against human experts with an average of 14 years of experience in fields ranging from finance to mechanical engineering. The latest AI News. Learn about LLMs, Gen AI and get ready for the rollout of AGI. Wes Roth covers the latest happenings in the world of OpenAI, Google, Anthropic, NVIDIA and Open Source AI. Emad Mostaque Interview "No One is Prepared" the next 1,000 days are CRUCIAL https://www.youtube.com/watch?v=07fuMWzFSUw OpenAI Introducing GPT-5.2 https://openai.com/index/introducing-gpt-5-2/ ______________________________________________ My Links 🔗 ➡️ Twitter: https://x.com/WesRothMoney ➡️ AI Newsletter: https://natural20.beehiiv.com/subscribe Want to work with me? Brand, sponsorship & business inquiries: [email protected] Check out my AI Podcast where me and Dylan interview AI experts: https://www.youtube.com/playlist?list=PLb1th0f6y4XSKLYenSVDUXFjSHsZTTfhk ______________________________________________ [00:00:00] Intro: 3D Spherical Conway's Game of Life [00:01:03] GPT 5.2 Release [00:02:55] Ethan Mollick & Noam Brown on GDP-Eval [00:03:52] Framer (Sponsor) [00:05:55] What is GDPval? [00:14:30] Economic Implications: When AI Outperforms Experts [00:16:15] Other Benchmarks: SWE-Bench, MATH, & ARC-AGI [00:18:10] Qualitative Leap: Cap Tables & Project Management [00:19:00] The Intelligence Curve: Performance vs. Compute Cost [00:20:40] 390x Cost Reduction in One Year [00:22:40] Addressing Skeptics: "Stochastic Parrots" vs. Real Utility [00:26:08] Model Testing #ai #openai #llm

Top Comments (10)

@TicketWhisperer 2025-12-12

Can you imagine, at one point we won't be smart enough to tell is something has gotten better in this space....

44 3 replies
@AIwithDwayne 2025-12-12

the "and wells fargo is on the list" comment killed me looool 10:10

27 1 replies
@theloniousMac 2025-12-12

People often criticize predictive language modeling by saying it doesn’t “know” answers and is therefore “just making things up.” That criticism misunderstands how both reasoning and knowledge actually work. When a computer answers “2 + 2 = 4,” it isn’t recalling a memorized fact in the human sense. It applies formal rules to arrive at a result. When a child who can’t yet read fluently sounds out a word letter by letter, they are not guessing randomly—they are applying learned patterns of language to reach the correct word. A skilled reader recognizes the word instantly, but the underlying goal is the same. Large language models operate in a similar way. They do not invent answers arbitrarily. They generate responses by applying learned statistical and structural patterns derived from vast amounts of real language, facts, and reasoning examples. The predictions are constrained by those patterns, just as arithmetic is constrained by mathematical rules and reading is constrained by phonetics and grammar. Prediction, in this context, does not mean randomness. It means selecting the most likely correct continuation based on structure, consistency, and prior knowledge embedded in the model. When an LLM gives a correct answer, it has followed a disciplined process grounded in learned relationships—not imagination. The key point is this: methods don’t matter—results do. Whether an answer is reached through calculation, phonetic decoding, memorization, or predictive modeling, the value lies in the correctness of the outcome. Prediction is not a weakness; it is simply another valid mechanism for reasoning toward truth. If the answer is right, the path taken to get there is irrelevant.

25 12 replies
@fleshtonegolem 2025-12-12

WOW the zip file result is CRAZY it's like a person sending you their work for approval.

12
@j2csharp 2025-12-12

"Wells Fargo is on the list." - I got that one, Wes.

10 1 replies
@WesRoth 2025-12-12

Launch your site for free at https://framer.link/WesRoth Use code WESROTH for a free month on Framer Pro. Build a site that looks hand-coded. Without hiring a developer.

7 3 replies
@sigigle 2025-12-12

21:35 Philosophy enthusiast here. In common parlance (no pun intended), "common" can suggest a majority occurrence, eg: "people commonly do X" can imply: "tend to more than not". So the phrase "not uncommon" is used to signal that an event is moderately frequent rather than rare, but avoids the suggestion that it might be "more often than not" that can be implied by the word common.

5
@djayjp 2025-12-12

Yep, I'm convinced GDPval is the most important benchmark around, in terms of imminent real-world impact.

3
@MakeTechPtyLtd 2025-12-12

What “shot” means in AI academia: In mainstream ML/LLM usage, a “shot” means an example provided in the model’s input context (the prompt), as a demonstration of the task. Zero-shot: no task examples in the prompt (just instructions, maybe constraints). One-shot: exactly one example (input → output) in the prompt. Few-shot: a small number of examples in the prompt (often 2–20ish; not a strict boundary).

2
@asdffdsa8863 2025-12-12

Wow great video Wes! Thanks for keeping tabs on all these new developments across all these models. It’s gotta be a full time job. Between the latest versions of Claude, Gemini, Grok, and now Chat, well they each have their specialties, but then that keeps changing every several weeks. Any way you can compile your personal top 3-10 models and what you like most of each? Of course, that runs stale after a month or two, but it would be very helpful to gain your insights. Happy Holidays!

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot