<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
        xmlns:video="http://www.google.com/schemas/sitemap-video/1.1">
  <url>
    <loc>https://www.pgelephant.com/ai/relevant-is-not-connected</loc>
    <lastmod>2026-09-30T14:03:34+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i4.ytimg.com/vi/wnto9FmzQjw/hqdefault.jpg</video:thumbnail_loc>
      <video:title>Relevant Is Not Connected</video:title>
      <video:description>Three relevant passages can still miss the path between them. An AI can hold the right documents, retrieve relevant chunks, and still miss the answer. Vector search finds similar meaning. It does not follow the link from vendor, to service, to product. Vector databases are not obsolete. Vector-only RAG is no longer enough.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/wnto9FmzQjw</video:player_loc>
      <video:publication_date>2026-09-30T14:03:34+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/vector-only-rag-is-dying-here-s-what-enterprise-ai-uses-now</loc>
    <lastmod>2026-09-30T08:54:29+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i1.ytimg.com/vi/822AZvyWeXM/hqdefault.jpg</video:thumbnail_loc>
      <video:title>Vector-Only RAG Is Dying. Here&apos;s What Enterprise AI Uses Now</video:title>
      <video:description>The right documents can still produce the wrong answer. Vector databases are not going away. Vector search alone does not answer every enterprise question. This video shows why the right documents can still produce the wrong answer, where vector search still wins, and what keyword search, knowledge graphs, reranking, and long context each do.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/822AZvyWeXM</video:player_loc>
      <video:publication_date>2026-09-30T08:54:29+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/can-you-spot-the-deepfake-the-eye-test-already-failed</loc>
    <lastmod>2026-09-22T09:24:16+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i1.ytimg.com/vi/4kGiDqJ4S24/hqdefault.jpg</video:thumbnail_loc>
      <video:title>Can You Spot the Deepfake? The Eye Test Already Failed</video:title>
      <video:description>Can you spot a deepfake? Blinking, teeth, and skin no longer reveal an AI face. Deepfake detection fails outside the lab. The generator already passed your eye test. The Model Already Passed Your Eye Test Why Deepfake Detectors Fail When They Leave the Lab | AI MECHANICS You will see: • GENERATOR - weaker models leaked.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/4kGiDqJ4S24</video:player_loc>
      <video:publication_date>2026-09-22T09:24:16+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/claude-mythos-found-10-000-serious-flaws-in-its-first-month</loc>
    <lastmod>2026-09-19T17:44:25+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i2.ytimg.com/vi/QkKKgKFZLsI/hqdefault.jpg</video:thumbnail_loc>
      <video:title>Claude Mythos Found 10,000 Serious Flaws in Its First Month</video:title>
      <video:description>This is not a chatbot. Anthropic built Claude Mythos to find security flaws humans had missed for decades. Then it restricted access. In 30 days, partners reported more than 10,000 high and critical findings. Claude Mythos Found 10,000 Serious Flaws in Its First Month Project Glasswing | AI MECHANICS Of the high and critical open-source subset reviewed by humans, 90.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/QkKKgKFZLsI</video:player_loc>
      <video:publication_date>2026-09-19T17:44:25+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/how-deepseek-shrinks-memory-and-openai-scales-thinking</loc>
    <lastmod>2026-09-17T09:38:54+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i4.ytimg.com/vi/wudmsqfVaJM/hqdefault.jpg</video:thumbnail_loc>
      <video:title>How DeepSeek Shrinks Memory and OpenAI Scales Thinking</video:title>
      <video:description>How DeepSeek Shrinks Memory and OpenAI Scales Thinking The Transformer Is Evolving | AI MECHANICS Every token your AI remembers costs GPU memory. At short context that looks harmless. At long context, the KV cache becomes one of inference’s biggest constraints. DeepSeek’s answer is not “the Transformer is dead. ” It is compress what you store, then spend compute where thinking helps.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/wudmsqfVaJM</video:player_loc>
      <video:publication_date>2026-09-17T09:38:54+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/gpt-6-astra-is-different-ai-now-finishes-the-job</loc>
    <lastmod>2026-09-09T16:48:28+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i2.ytimg.com/vi/EbxLtdJgFnc/hqdefault.jpg</video:thumbnail_loc>
      <video:title>GPT-6 Astra Is Different: AI Now Finishes the Job</video:title>
      <video:description>Most models stop at an answer. GPT-6 Astra is built to finish a job. This AI Mechanics episode explains the real shift: from one-turn chat to goal, reason, act, observe, adapt, complete. Computer use, agentic coding, research loops, mid-turn steering, async tools, and a token price about 2. 5 times Sol. No AGI claims. Vendor benchmarks are labeled OpenAI-reported.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/EbxLtdJgFnc</video:player_loc>
      <video:publication_date>2026-09-09T16:48:28+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/kimi-k3-2-8t-parameters-but-only-104b-active-per-token</loc>
    <lastmod>2026-09-03T04:04:03+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i2.ytimg.com/vi/UeoNl5U_f5E/hqdefault.jpg</video:thumbnail_loc>
      <video:title>Kimi K3: 2.8T Parameters, But Only 104B Active Per Token</video:title>
      <video:description>Kimi K3 has 2. 8 trillion parameters-but only about 104 billion are activated for each token. This visual deep dive explains how 16 of 896 routed experts, 2 shared experts, Stable LatentMoE, and Kimi Delta Attention make that possible. Kimi K3 is a Mixture-of-Experts model with a 1M-token context window. We follow one token through expert routing, distributed communication, the 7168 → 3584 latent projection, Quantile Balancing, KDA, Gated MLA, and Block Attention Residuals.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/UeoNl5U_f5E</video:player_loc>
      <video:publication_date>2026-09-03T04:04:03+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/how-chatgpt-actually-writes-next-token-prediction-explained</loc>
    <lastmod>2026-08-31T14:53:00+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i2.ytimg.com/vi/-mOKLAP5ct4/hqdefault.jpg</video:thumbnail_loc>
      <video:title>How ChatGPT Actually Writes: Next Token Prediction Explained</video:title>
      <video:description>What is an LLM? ChatGPT generates answers by predicting the next token. Beginner visual guide-no math required. Tokens, weights, context, and hallucinations. What you will learn • What LLM means • How next-token prediction builds an answer • Where probabilities come from • Context vs memory • LLM vs the full ChatGPT-style app Playlists Beginner AI: https://www.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/-mOKLAP5ct4</video:player_loc>
      <video:publication_date>2026-08-31T14:53:00+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/why-ai-models-think-longer-before-answering-test-time-compute-explained</loc>
    <lastmod>2026-08-30T11:12:31+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i1.ytimg.com/vi/td3bOBa7KMY/hqdefault.jpg</video:thumbnail_loc>
      <video:title>Why AI Models Think Longer Before Answering | Test-Time Compute Explained</video:title>
      <video:description>Why do reasoning models spend more compute before answering? Test-time compute. Best-of-N, backtracking, verification, rewards, DeepSeek-R1, and tool use-visually explained. What you will learn • Training-time vs test-time scaling • Reasoning budgets and candidate paths • Best-of-N, backtracking, and verification • Outcome vs process rewards • When thinking longer helps-and when it fails Playlists DEEP AI: https://www.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/td3bOBa7KMY</video:player_loc>
      <video:publication_date>2026-08-30T11:12:31+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/speculative-decoding-faster-llm-token-generation</loc>
    <lastmod>2026-08-30T10:34:14+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i4.ytimg.com/vi/oLazZ0I8-aQ/hqdefault.jpg</video:thumbnail_loc>
      <video:title>Speculative Decoding: Faster LLM Token Generation</video:title>
      <video:description>What is speculative decoding? A small draft model proposes tokens; the large LLM verifies them. Same output distribution as the big model-faster when the draft is often right. What you will learn • Why one token per forward pass wastes GPU parallelism • Draft model vs target model • One-pass verification and rejection • Why draft quality controls the speedup • How it sits next to GQA, quantization, and PagedAttention Playlists DEEP AI: https://www.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/oLazZ0I8-aQ</video:player_loc>
      <video:publication_date>2026-08-30T10:34:14+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/why-llm-prediction-is-compression-entropy-explained</loc>
    <lastmod>2026-08-28T09:05:12+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i1.ytimg.com/vi/T_u6Oyyp1ag/hqdefault.jpg</video:thumbnail_loc>
      <video:title>Why LLM Prediction Is Compression (Entropy Explained)</video:title>
      <video:description>What is entropy? Next-token prediction and compression are the same math. From Shannon’s information theory to why language models train with cross-entropy. Watch next: Transformer Attention Explained - https://www. youtube. com/watch? v=rbrSteyXx_0 What you will learn • Variable-length and prefix-free codes • Self-information I = -log₂ p • Why prediction equals compression • Entropy rate of language • How this leads to LLM training loss Playlists LLM Fundamentals: https://www.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/T_u6Oyyp1ag</video:player_loc>
      <video:publication_date>2026-08-28T09:05:12+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/how-ai-agents-work-tools-memory-mcp-guardrails</loc>
    <lastmod>2026-08-27T20:08:03+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i4.ytimg.com/vi/kY--WQxI2Rw/hqdefault.jpg</video:thumbnail_loc>
      <video:title>How AI Agents Work: Tools, Memory, MCP, Guardrails</video:title>
      <video:description>How AI agents work beyond the LLM: prompt → tool call → policy → runtime → observation. Follow one agent from a user request to a controlled real-world action-plus memory, MCP, and guardrails. What you will learn • Why writing a plan is not taking an action • Structured tool calls and the agent loop • Session context vs durable memory • MCP discovery is not permission • Policy gates, human approval, budgets, and tracing Playlists AI Agents: https://www.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/kY--WQxI2Rw</video:player_loc>
      <video:publication_date>2026-08-27T20:08:03+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
  <url>
    <loc>https://www.pgelephant.com/ai/how-vector-databases-work-embeddings-hnsw-semantic-search</loc>
    <lastmod>2026-08-25T22:59:45+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.85</priority>
    <video:video>
      <video:thumbnail_loc>https://i1.ytimg.com/vi/4eCj3GOhpW4/hqdefault.jpg</video:thumbnail_loc>
      <video:title>How Vector Databases Work: Embeddings, HNSW &amp; Semantic Search</video:title>
      <video:description>One query. A million stored vectors. Do we compare against every one? That is exact nearest neighbor, and it is correct, and it is linear. At that scale, latency is the product. A point in space is not a search. This lecture is the store those numbers live in, and the hop that finds the nearest neighbor without reading every row.</video:description>
      <video:player_loc allow_embed="yes">https://www.youtube.com/embed/4eCj3GOhpW4</video:player_loc>
      <video:publication_date>2026-08-25T22:59:45+00:00</video:publication_date>
      <video:uploader info="https://www.youtube.com/@DrIbrarAhmedAI">Dr. Ibrar Ahmed</video:uploader>
      <video:family_friendly>yes</video:family_friendly>
      <video:requires_subscription>no</video:requires_subscription>
      <video:live>no</video:live>
    </video:video>
  </url>
</urlset>