GPT 5.2 is the first HUMAN LABOR replacement
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
Mythos 5 is WILD...
Wes Roth
59.1k views
Hermes Agent is INSANE...
Wes Roth
37.3k views
OpenAI's GPT 5.5 is wild...
Wes Roth
61.0k views
HERMES AGENT SETUP: the OpenClaw killer is here
Wes Roth
25.5k views
GEMINI 3.1 PRO is the new era...
Wes Roth
40.8k views
GROK 4.20 is... different
Wes Roth
36.8k views
White Replacement Is Real
StevenCrowder
11.5k views
we just arrived at the "WTF" moment in AI
Wes Roth
62.8k views
the world wasn't ready for Gemini 3
Wes Roth
42.2k views
Keanu Reeves revisits "The Replacements" | New Heights Film Club
New Heights
510.9k views
Top Comments (10)
Can you imagine, at one point we won't be smart enough to tell is something has gotten better in this space....
the "and wells fargo is on the list" comment killed me looool 10:10
People often criticize predictive language modeling by saying it doesn’t “know” answers and is therefore “just making things up.” That criticism misunderstands how both reasoning and knowledge actually work. When a computer answers “2 + 2 = 4,” it isn’t recalling a memorized fact in the human sense. It applies formal rules to arrive at a result. When a child who can’t yet read fluently sounds out a word letter by letter, they are not guessing randomly—they are applying learned patterns of language to reach the correct word. A skilled reader recognizes the word instantly, but the underlying goal is the same. Large language models operate in a similar way. They do not invent answers arbitrarily. They generate responses by applying learned statistical and structural patterns derived from vast amounts of real language, facts, and reasoning examples. The predictions are constrained by those patterns, just as arithmetic is constrained by mathematical rules and reading is constrained by phonetics and grammar. Prediction, in this context, does not mean randomness. It means selecting the most likely correct continuation based on structure, consistency, and prior knowledge embedded in the model. When an LLM gives a correct answer, it has followed a disciplined process grounded in learned relationships—not imagination. The key point is this: methods don’t matter—results do. Whether an answer is reached through calculation, phonetic decoding, memorization, or predictive modeling, the value lies in the correctness of the outcome. Prediction is not a weakness; it is simply another valid mechanism for reasoning toward truth. If the answer is right, the path taken to get there is irrelevant.
WOW the zip file result is CRAZY it's like a person sending you their work for approval.
"Wells Fargo is on the list." - I got that one, Wes.
Launch your site for free at https://framer.link/WesRoth Use code WESROTH for a free month on Framer Pro. Build a site that looks hand-coded. Without hiring a developer.
21:35 Philosophy enthusiast here. In common parlance (no pun intended), "common" can suggest a majority occurrence, eg: "people commonly do X" can imply: "tend to more than not". So the phrase "not uncommon" is used to signal that an event is moderately frequent rather than rare, but avoids the suggestion that it might be "more often than not" that can be implied by the word common.
Yep, I'm convinced GDPval is the most important benchmark around, in terms of imminent real-world impact.
What “shot” means in AI academia: In mainstream ML/LLM usage, a “shot” means an example provided in the model’s input context (the prompt), as a demonstration of the task. Zero-shot: no task examples in the prompt (just instructions, maybe constraints). One-shot: exactly one example (input → output) in the prompt. Few-shot: a small number of examples in the prompt (often 2–20ish; not a strict boundary).
Wow great video Wes! Thanks for keeping tabs on all these new developments across all these models. It’s gotta be a full time job. Between the latest versions of Claude, Gemini, Grok, and now Chat, well they each have their specialties, but then that keeps changing every several weeks. Any way you can compile your personal top 3-10 models and what you like most of each? Of course, that runs stale after a month or two, but it would be very helpful to gain your insights. Happy Holidays!
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
Can you imagine, at one point we won't be smart enough to tell is something has gotten better in this space....
the "and wells fargo is on the list" comment killed me looool 10:10
People often criticize predictive language modeling by saying it doesn’t “know” answers and is therefore “just making things up.” That criticism misunderstands how both reasoning and knowledge actually work. When a computer answers “2 + 2 = 4,” it isn’t recalling a memorized fact in the human sense. It applies formal rules to arrive at a result. When a child who can’t yet read fluently sounds out a word letter by letter, they are not guessing randomly—they are applying learned patterns of language to reach the correct word. A skilled reader recognizes the word instantly, but the underlying goal is the same. Large language models operate in a similar way. They do not invent answers arbitrarily. They generate responses by applying learned statistical and structural patterns derived from vast amounts of real language, facts, and reasoning examples. The predictions are constrained by those patterns, just as arithmetic is constrained by mathematical rules and reading is constrained by phonetics and grammar. Prediction, in this context, does not mean randomness. It means selecting the most likely correct continuation based on structure, consistency, and prior knowledge embedded in the model. When an LLM gives a correct answer, it has followed a disciplined process grounded in learned relationships—not imagination. The key point is this: methods don’t matter—results do. Whether an answer is reached through calculation, phonetic decoding, memorization, or predictive modeling, the value lies in the correctness of the outcome. Prediction is not a weakness; it is simply another valid mechanism for reasoning toward truth. If the answer is right, the path taken to get there is irrelevant.
WOW the zip file result is CRAZY it's like a person sending you their work for approval.
"Wells Fargo is on the list." - I got that one, Wes.
Launch your site for free at https://framer.link/WesRoth Use code WESROTH for a free month on Framer Pro. Build a site that looks hand-coded. Without hiring a developer.
21:35 Philosophy enthusiast here. In common parlance (no pun intended), "common" can suggest a majority occurrence, eg: "people commonly do X" can imply: "tend to more than not". So the phrase "not uncommon" is used to signal that an event is moderately frequent rather than rare, but avoids the suggestion that it might be "more often than not" that can be implied by the word common.
Yep, I'm convinced GDPval is the most important benchmark around, in terms of imminent real-world impact.
What “shot” means in AI academia: In mainstream ML/LLM usage, a “shot” means an example provided in the model’s input context (the prompt), as a demonstration of the task. Zero-shot: no task examples in the prompt (just instructions, maybe constraints). One-shot: exactly one example (input → output) in the prompt. Few-shot: a small number of examples in the prompt (often 2–20ish; not a strict boundary).
Wow great video Wes! Thanks for keeping tabs on all these new developments across all these models. It’s gotta be a full time job. Between the latest versions of Claude, Gemini, Grok, and now Chat, well they each have their specialties, but then that keeps changing every several weeks. Any way you can compile your personal top 3-10 models and what you like most of each? Of course, that runs stale after a month or two, but it would be very helpful to gain your insights. Happy Holidays!