{
  "version": 1,
  "language": "vi",
  "items": [
    {
      "id": 1,
      "text": "Bạn đang ngồi trước bàn làm việc, mắt vừa rời khỏi màn hình sau một ngày dài, và bỗng chiếc ổ cắm dưới chân tường nhìn lên với vẻ ngạc nhiên: hai lỗ tròn như đôi mắt, khe nhỏ phía dưới như một cái miệng đang há ra. Ngoài đường, cụm đèn trước của một chiếc xe có vẻ cau có. Mặt tiền căn nhà đối diện trông như đang buồn ngủ. Trong cốc cà phê, một vệt bọt mỏng bất ngờ tạo thành gương mặt đang mỉm cười. Bạn biết rất rõ không có ai ở đó. Vật thể không thật sự có khuôn mặt, nhưng não bạn đang hoàn thành một giả thuyết.",
      "start": 0.0,
      "end": 28.36,
      "duration": 28.36,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Information-dense educational stick-figure tableau. Foreground: the recurring white round-headed protagonist in an orange shirt swivels away from a laptop and points in surprise at a wall outlet whose two sockets and lower slot resemble a face. Midground: a four-panel sequence shows angry-looking car headlights and grille, a sleepy house facade with window-eyes, smiling coffee foam, and the protagonist correctly recognizing each as an object. Background: a workspace wall continues into a street and cafe context, with faint neural lines linking all four accidental faces to a brain icon. Include at least four narration details: tired eyes leaving the screen, surprised outlet, scowling car, sleepy house, smiling coffee foam, and awareness that no real person is present. Explanatory device: arrows connect pairs of eye-like points plus a mouth-like line to a brain hypothesis bubble containing a generic face template. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. STRICT PROTAGONIST FACE OVERRIDE: every appearance of the orange-shirt protagonist has exactly two tiny solid black circular dot eyes, two separate thick black eyebrow strokes above them, and one minimal black line mouth. No eyelids, no sclera, no pupils, no large eyes, no anime eyes, no half-closed eyes, no eye whites. Object pareidolia faces may use slots and shapes, but the protagonist must keep the exact dot-eye design. FINAL SOLID-FILL OVERRIDE: simplify each panel into large clean vector-like shapes. Use one exact solid color per wall, floor, card, car body, windshield, house wall, roof, coffee cup, desk, brain, and background. Eliminate tiny decorative texture, highlight patches, reflections, tonal variations, mottling, and micro-details. Preserve all required objects, arrows, face-template flow, and the exact dot-eye protagonist. Crisp black outlines and uniform fills only.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_001.png"
    },
    {
      "id": 2,
      "text": "Hiện tượng này được gọi là pareidolia, thường được dịch là ảo giác nhận dạng hình mẫu, hay cụ thể hơn trong trường hợp này là pareidolia khuôn mặt. Não tìm thấy một hình ảnh quen thuộc và có ý nghĩa trong những tín hiệu vốn mơ hồ hoặc ngẫu nhiên: khuôn mặt trên mây, con vật trong vết nứt tường, bóng người trong tấm rèm, giọng nói trong tiếng nhiễu. Đây không nhất thiết là lỗi nhận thức, càng không tự động có nghĩa rằng bạn đang mất liên hệ với thực tế. Phần lớn thời gian, nó chỉ cho thấy hệ thống nhận dạng của não hoạt động nhanh, chủ động và hơi thiên lệch về phía an toàn.",
      "start": 28.36,
      "end": 59.12,
      "duration": 30.76,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Educational stick-figure classification board with the recurring white round-headed protagonist in an orange shirt examining ambiguous patterns through a magnifier. Foreground: the protagonist sorts cards showing a face in clouds, an animal in a wall crack, a person-like curtain shadow, and a waveform hidden in static. Midground: each real stimulus remains visible underneath a translucent interpretation overlay, emphasizing that meaningful patterns arise from ambiguous or random signals. Background: a brain-shaped detection system rapidly scans noisy dots, with a green reality-check shield indicating preserved contact with reality and a slightly lowered decision threshold. Include at least four narration details: cloud face, animal-shaped crack, human silhouette in a curtain, voice-like pattern in noise, normal recognition bias, and safety-oriented detection. Explanatory device: split-screen labeled conceptually by icons only, contrasting raw stimulus on the left with perceived pattern on the right, joined by inference arrows. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Pareidolia\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Pareidolia"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_002.png"
    },
    {
      "id": 3,
      "text": "Khuôn mặt là một trong những tín hiệu quan trọng nhất trong đời sống xã hội. Từ khuôn mặt, bạn ước lượng một người đang chú ý đến đâu, có thân thiện không, có tức giận không, có quen thuộc không, và có thể sắp làm gì. Trước khi hiểu lời nói, trẻ nhỏ đã bị thu hút bởi những cấu hình giống khuôn mặt. Trong suốt cuộc đời, hàng nghìn quyết định xã hội phụ thuộc vào khả năng phát hiện và đọc khuôn mặt gần như tức thì. Vì vậy, não không chờ một bức chân dung hoàn hảo mới bắt đầu xử lý. Nó có thể khởi động chỉ từ vài dấu hiệu thô.",
      "start": 59.12,
      "end": 88.04,
      "duration": 28.92,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Information-rich social perception diagram in a clean stick-figure educational style. Foreground: the recurring white round-headed protagonist in an orange shirt reads four cues from another stick figure's face—gaze direction, friendly smile, angry brows, and familiarity—while action arrows predict where that person may move next. Midground: an infant turns toward a simple two-eyes-and-mouth configuration before nearby speech bubbles appear, and an adult crowd demonstrates rapid face detection during everyday decisions. Background: a lifetime ribbon runs from infancy to adulthood with thousands of tiny social encounters, while a coarse face template activates before a detailed portrait is complete. Include at least four narration details: attention direction, friendliness, anger, familiarity, likely next action, infant attraction, and lifelong instant decisions. Explanatory device: progressive diagram from three crude marks to a complete face, with arrows showing processing begins before perfect visual detail arrives. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_003.png"
    },
    {
      "id": 4,
      "text": "Carl Sagan từng phổ biến một cách giải thích giàu sức gợi trong cuốn “The Demon-Haunted World”: khả năng nhận ra khuôn mặt nhanh chóng có thể mang giá trị sống còn, đặc biệt với trẻ sơ sinh phụ thuộc vào người chăm sóc. Đây là một cách diễn giải phổ thông có ảnh hưởng, không phải tự thân là bằng chứng thực nghiệm cho toàn bộ cơ chế. Nhưng nó nắm được một ý quan trọng: đối với hệ thần kinh, bỏ sót một khuôn mặt thật đôi khi có thể đắt giá hơn nhìn nhầm một khuôn mặt không tồn tại.",
      "start": 88.04,
      "end": 113.84,
      "duration": 25.8,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Educational evolutionary tradeoff scene centered on the recurring white round-headed protagonist in an orange shirt holding an open Carl Sagan book while carefully separating an influential explanation from direct experimental proof. Foreground: beside the protagonist, an infant rapidly detects a caregiver's face emerging from visual clutter. Midground: two possible outcomes show detecting the caregiver early versus overlooking the caregiver, with the latter carrying a larger warning symbol. Background: a simplified ancestral environment contains partially hidden allies, strangers, and predators, illustrating why rapid face detection could have survival value. Include at least four narration details: Carl Sagan's popular explanation, infant dependence on caregivers, possible survival value, the distinction between interpretation and full empirical evidence, and the higher cost of missing a real face. Explanatory device: balance scale comparing one small false-alarm cost with one large missed-face cost, plus a caution callout separating hypothesis from evidence. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_004.png"
    },
    {
      "id": 5,
      "text": "Hãy hình dung hai loại sai lầm. Sai lầm thứ nhất là false positive: bạn tưởng bụi cây có người, nhưng thật ra chỉ là bóng lá. Sai lầm thứ hai là false negative: có một người đang nhìn hoặc một mối đe dọa đang tiến đến, nhưng bạn không phát hiện. Trong nhiều hoàn cảnh xã hội và tiến hóa, false positive chỉ khiến bạn giật mình trong chốc lát; false negative có thể khiến bạn bỏ lỡ người chăm sóc, đồng minh, đối thủ, kẻ săn mồi hoặc một tín hiệu cảm xúc quan trọng. Một hệ thống được điều chỉnh để báo động sớm vì thế có thể chấp nhận nhận nhầm đôi chút.",
      "start": 113.84,
      "end": 132.66,
      "duration": 18.82,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Dense split-screen risk comparison in an educational stick-figure style. Foreground: the recurring white round-headed protagonist in an orange shirt reacts on the left to a bush shadow mistakenly perceived as a person, then on the right fails to notice a real approaching figure. Midground: the false-positive side resolves harmlessly into leaves after a brief startle; the false-negative side branches into missed caregiver, ally, rival, predator, and emotional-warning icons. Background: an ancestral-social landscape visualizes unequal consequences, with a small cost meter on the false alarm and a large cost meter on the missed detection. Include at least four narration details: bush mistaken for a person, brief startle, real observer or threat missed, missed caregiver or ally, missed opponent or predator, and lost emotional signal. Explanatory device: explicit two-column decision matrix with outcome arrows and weighted consequence scales. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Báo nhầm\", \"Bỏ sót\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Báo nhầm",
        "Bỏ sót"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_005.png"
    },
    {
      "id": 6,
      "text": "Face-detection bias, hay thiên lệch phát hiện khuôn mặt, không có nghĩa não lúc nào cũng sai. Nó có nghĩa ngưỡng để cân nhắc giả thuyết “khuôn mặt” được đặt khá thấp.",
      "start": 132.66,
      "end": 142.02,
      "duration": 9.36,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Compact but information-dense detector-threshold infographic. Foreground: the recurring white round-headed protagonist in an orange shirt adjusts a brain-shaped face detector whose slider sits near a low evidence threshold, while several ambiguous three-mark patterns enter the scanner. Midground: genuine faces and a few face-like objects both cross the alert line, showing sensitivity without implying constant error. Background: a probability gauge displays weak-to-strong face evidence, with a highlighted zone where the brain begins considering the face hypothesis before certainty. Include at least four narration details: face-detection bias, a low threshold for considering a face, correct detections still occurring, occasional object false alarms, and hypothesis generation rather than final belief. Explanatory device: threshold graph with incoming evidence dots, a horizontal decision line, and arrows showing which signals trigger consideration. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Ngưỡng phát hiện\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Ngưỡng phát hiện"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_006.png"
    },
    {
      "id": 7,
      "text": "Điều đáng chú ý là thị giác không vận hành như máy ảnh chỉ ghi lại thế giới rồi chuyển nguyên vẹn vào đầu. Ánh sáng đi vào mắt tạo nên dữ liệu chưa đầy đủ, có nhiễu và thường mơ hồ. Não phải liên tục suy luận nguyên nhân nào ngoài thế giới có khả năng tạo ra dữ liệu đó. Cách nhìn này thường được mô tả bằng các khái niệm top-down processing và predictive processing. Tín hiệu từ mắt đi lên cung cấp bằng chứng, trong khi kinh nghiệm, kỳ vọng, bối cảnh và mục tiêu chú ý đi xuống để định hình cách bằng chứng được giải thích.",
      "start": 142.02,
      "end": 170.7,
      "duration": 28.68,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Educational predictive-processing systems diagram with the recurring white round-headed protagonist in an orange shirt observing a dim, partly occluded object. Foreground: photons enter the protagonist's eyes as incomplete, noisy fragments rather than a finished picture. Midground: upward arrows carry sensory evidence toward a brain, while downward arrows carry experience, expectation, context, and attention toward the same interpretation workspace. Background: several possible world causes—face, outlet, mask, and random pattern—compete until the brain selects the most probable explanation. Include at least four narration details: vision is not a passive camera, retinal data are incomplete, noise creates ambiguity, the brain infers external causes, sensory evidence travels upward, and expectations and goals shape interpretation downward. Explanatory device: bidirectional flowchart contrasting bottom-up evidence with top-down prediction and a central hypothesis-comparison panel. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_007.png"
    },
    {
      "id": 8,
      "text": "Vì vậy, pareidolia dễ mạnh hơn khi thông tin thị giác kém rõ. Trong tầm nhìn ngoại vi, độ phân giải chi tiết thấp hơn vùng trung tâm mà bạn đang nhìn trực tiếp. Một vật ở khóe mắt có thể trở thành bóng người; khi quay đầu nhìn thẳng, nó lại chỉ là chiếc áo khoác. Ở khoảng cách xa, các đường nét nhỏ bị mất đi và chỉ còn cấu trúc tổng quát. Trong ảnh nén, hình thu nhỏ, camera thiếu sáng hoặc màn hình có nhiễu, não nhận được ít bằng chứng từ dưới lên hơn, nên kỳ vọng từ trên xuống có nhiều khoảng trống để lấp đầy.",
      "start": 170.7,
      "end": 182.38,
      "duration": 11.68,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Multi-condition visual ambiguity laboratory in an educational stick-figure style. Foreground: the recurring white round-headed protagonist in an orange shirt first sees a person-like coat rack from the corner of the eye, then turns to inspect it directly and recognizes a coat. Midground: three comparison stations show a distant object losing fine contours, a compressed thumbnail retaining only broad face layout, and a low-light noisy camera frame generating false features. Background: a retina map contrasts sharp central vision with low-detail peripheral vision, while a top-down prediction cloud expands wherever bottom-up evidence becomes weak. Include at least four narration details: peripheral low resolution, coat becoming a shadow person, direct gaze restoring object identity, distance removing detail, compression and thumbnails, low light and screen noise. Explanatory device: split-screen before-and-after gaze shift plus a resolution gradient and arrows showing prediction filling missing information. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Ngoại vi\", \"Nhìn thẳng\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Ngoại vi",
        "Nhìn thẳng"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_008.png"
    },
    {
      "id": 9,
      "text": "Các nghiên cứu của nhà tâm lý học phát triển Kang Lee và cộng sự đã khảo sát trực tiếp việc não phản ứng với những vật thể có bố cục giống khuôn mặt. Trong một nghiên cứu dùng fMRI do Jiangang Liu, Jun Li, Lu Feng, Ling Li, Jie Tian và Kang Lee công bố, những hình ảnh pareidolia khuôn mặt được so sánh với các vật thể tương tự nhưng không gợi mặt rõ ràng. Kết quả cho thấy trải nghiệm nhìn thấy khuôn mặt trong vật vô tri không chỉ là cách nói vui sau khi quan sát.",
      "start": 182.38,
      "end": 210.02,
      "duration": 27.64,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Research-method scene featuring Kang Lee's team in a clear educational stick-figure laboratory. Foreground: the recurring white round-headed protagonist in an orange shirt lies in an fMRI scanner while viewing matched image pairs: an object with a face-like layout and a visually similar object without a convincing face. Midground: stick-figure researchers Kang Lee, Jiangang Liu, Jun Li, Lu Feng, Ling Li, and Jie Tian compare scanner outputs and behavioral face reports on a large console. Background: controlled stimulus cards preserve object category and visual complexity while varying facial configuration, preventing the result from looking like a casual joke or anecdote. Include at least four narration details: Kang Lee's developmental psychology research, direct study of face-like objects, fMRI methodology, matched pareidolia and non-face objects, a multi-researcher study, and measurable neural involvement. Explanatory device: experimental pipeline diagram from stimulus pair to participant to scanner to comparison graph. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_009.png"
    },
    {
      "id": 10,
      "text": "Các vùng thuộc hệ thống xử lý khuôn mặt, đặc biệt vùng thường được gọi là fusiform face area, có thể tham gia khi cấu hình mơ hồ được diễn giải thành mặt.",
      "start": 210.02,
      "end": 218.1,
      "duration": 8.08,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Focused brain-activation explainer in an information-dense stick-figure style. Foreground: the recurring white round-headed protagonist in an orange shirt views an ambiguous outlet-like arrangement that alternates between object interpretation and face interpretation. Midground: a side-view brain cutaway highlights the ventral temporal face-processing region more strongly when the same configuration is interpreted as a face. Background: a small set of comparison cards includes a real face, a face-like object, and an ordinary object, each feeding different activation bars. Include at least four narration details: ambiguous configuration, interpretation as a face, participation of the face-processing system, fusiform region involvement, comparison with ordinary objects, and activation without claiming the object is alive. Explanatory device: arrow-linked stimulus-to-percept-to-brain activation diagram with color-coded response bars. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"FFA\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "FFA"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_010.png"
    },
    {
      "id": 11,
      "text": "Fusiform face area, viết tắt là FFA, nằm trong vùng hồi hình thoi ở mặt dưới thùy thái dương. Công trình nổi tiếng của Nancy Kanwisher, Josh McDermott và Marvin Chun vào cuối thập niên 1990 cho thấy một vùng tại đây phản ứng mạnh hơn với khuôn mặt so với nhiều loại vật thể khác trong các thí nghiệm fMRI. Tên gọi FFA rất tiện, nhưng không nên hiểu nó như một chiếc hộp duy nhất chỉ có nhiệm vụ bật lên mỗi khi có mặt.",
      "start": 218.1,
      "end": 244.62,
      "duration": 26.52,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Anatomical and historical educational infographic. Foreground: the recurring white round-headed protagonist in an orange shirt points to the underside of a rotating brain model, where the fusiform gyrus in the ventral temporal lobe is highlighted. Midground: stick-figure researchers Nancy Kanwisher, Josh McDermott, and Marvin Chun stand beside a late-1990s fMRI comparison showing stronger responses to faces than to several object categories. Background: a network of connected brain regions surrounds the highlighted FFA so it is not depicted as a single isolated face box. Include at least four narration details: full name fusiform face area, FFA abbreviation, fusiform gyrus location, underside of temporal lobe, late-1990s research, stronger face response, and caution against one-box simplification. Explanatory device: labeled anatomical inset, historical timeline marker, and network arrows extending beyond FFA. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"FFA\", \"Hồi hình thoi\", \"Thập niên 1990\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. SCIENTIFIC AND COMPOSITION CORRECTION: Show FFA as one labeled component inside a clearly connected face-processing network, not a lone master hub. Label the fusiform gyrus and connect FFA to distinct icon-only nodes for visual face structure, dynamic social cues, emotion/salience, and identity memory. Remove the crossed-out face and dangling unexplained arrows. Keep the response chart comparing faces with ordinary objects. FINAL TEXT OVERRIDE: the only readable words anywhere are exactly [\"FFA\", \"Hồi hình thoi\", \"Mạng xử lý khuôn mặt\"]; render each exactly once with correct Vietnamese diacritics; no other text, English, numbers, pseudo-text, or watermark. FINAL REGENERATION OVERRIDE: Simplify the right-side network to exactly five closed circular nodes: central FFA plus four surrounding icon-only nodes for face structure, dynamic gaze, emotion/salience, and identity memory. Every arrow must start and end at a node boundary. No crossing arrows, no dangling lines, no crossed-out face. Keep only the exact permitted labels already specified. ABSOLUTE TEXT COUNT OVERRIDE: render the string FFA exactly once total, only inside the central network node. The anatomical brain callout uses color and an unlabeled pointer only. Render Hồi hình thoi exactly once and Mạng xử lý khuôn mặt exactly once. No duplicated labels or other characters.",
      "visible_text": [
        "FFA",
        "Hồi hình thoi",
        "Mạng xử lý khuôn mặt"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_011.png"
    },
    {
      "id": 12,
      "text": "Nhận dạng khuôn mặt phụ thuộc vào cả một mạng lưới, và vẫn có tranh luận khoa học về mức độ chuyên biệt của từng vùng, vai trò của kinh nghiệm, cũng như cách biểu diễn khuôn mặt được phân bố.",
      "start": 244.62,
      "end": 254.46,
      "duration": 9.84,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Scientific-debate network map in a rich educational stick-figure style. Foreground: the recurring white round-headed protagonist in an orange shirt connects multiple brain nodes rather than selecting one central box. Midground: three researcher groups compare competing models—highly specialized regions, expertise shaped by experience, and distributed face representations—using different diagrams on transparent boards. Background: sensory, memory, attention, identity, emotion, and motion nodes form an interconnected face-recognition network, with dotted lines marking unresolved relationships. Include at least four narration details: recognition depends on a network, specialization remains debated, experience may shape responses, representations may be distributed, multiple regions cooperate, and scientific uncertainty is preserved. Explanatory device: three-way model comparison with question marks, bidirectional arrows, and a shared evidence panel. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Chuyên biệt?\", \"Kinh nghiệm?\", \"Phân tán?\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Chuyên biệt?",
        "Kinh nghiệm?",
        "Phân tán?"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_012.png"
    },
    {
      "id": 13,
      "text": "Nouchine Hadjikhani và các cộng sự từng báo cáo rằng những hình ảnh pareidolia có thể huy động các vùng thuộc mạng lưới nhận diện khuôn mặt, kể cả khi người xem biết đó chỉ là đồ vật. Điều này giúp giải thích một trải nghiệm quen thuộc: kiến thức “đây là ổ cắm” và ấn tượng “nó giống một gương mặt” có thể cùng tồn tại. Nhận thức không buộc phải chọn duy nhất một nhãn. Một mức xử lý nhận ra vật thể thật; mức khác phát hiện cấu hình xã hội nổi bật được đặt lên trên nó.",
      "start": 254.46,
      "end": 275.92,
      "duration": 21.46,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Dual-perception demonstration referencing Nouchine Hadjikhani's research. Foreground: the recurring white round-headed protagonist in an orange shirt looks at a wall outlet and holds two simultaneous thought cards: one accurately identifies an electrical object, while the other detects a surprised facial configuration. Midground: a layered brain diagram shows object-recognition processing and face-network activity operating in parallel rather than forcing one label to replace the other. Background: Hadjikhani and colleagues review face-like object scans, while additional examples—a kettle, house facade, and car front—show the same coexistence of knowledge and impression. Include at least four narration details: Hadjikhani's findings, recruitment of the face network, conscious knowledge that the stimulus is an object, simultaneous face impression, multiple processing levels, and a socially salient overlay. Explanatory device: transparent double-layer overlay and two parallel arrows converging on one conscious scene. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Ổ cắm\", \"Giống khuôn mặt\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Ổ cắm",
        "Giống khuôn mặt"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_013.png"
    },
    {
      "id": 14,
      "text": "Nghiên cứu bằng EEG và MEG cũng cung cấp góc nhìn về thời gian. Các đáp ứng điện não liên quan đến xử lý khuôn mặt, thường được nghiên cứu qua thành phần N170 xuất hiện khoảng hơn một phần mười giây sau khi hình ảnh hiện ra, có thể bị ảnh hưởng bởi những hình vô tri trông giống mặt. Không phải mọi nghiên cứu đều cho kết quả giống hệt nhau, và cách thiết kế kích thích rất quan trọng. Tuy nhiên, bức tranh chung cho thấy pareidolia có thể đi vào những giai đoạn xử lý thị giác tương đối sớm, chứ không chỉ xuất hiện sau một chuỗi suy luận chậm bằng ngôn ngữ.",
      "start": 275.92,
      "end": 307.26,
      "duration": 31.34,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Time-resolved neuroscience laboratory with dense but readable educational graphics. Foreground: the recurring white round-headed protagonist in an orange shirt wears an EEG cap while face-like object images flash on a monitor; a nearby MEG helmet provides a second timing method. Midground: an electrical waveform highlights the N170 response shortly after stimulus onset, with comparison traces for real faces, pareidolia images, and ordinary objects. Background: multiple experiment cards vary stimulus design and produce not-quite-identical results, showing why methodology matters and conclusions require caution. Include at least four narration details: EEG, MEG, face-related electrical responses, N170, onset a little over one tenth of a second after viewing, influence of face-like inanimate images, and variation across studies. Explanatory device: millisecond timeline from image onset to early visual activity to N170, plus overlaid waveform comparison. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"EEG / MEG\", \"N170\", \"≈ 170 ms\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. SCIENTIFIC AND COMPOSITION CORRECTION: Create two clearly separate parallel branches. Left branch: EEG cap -> EEG waveform labeled N170. Right branch: MEG helmet scanner -> MEG waveform labeled M170. Both converge only on a timing marker labeled approximately 170 ms. Never label an MEG signal N170 and never depict a hybrid EEG-MEG device. FINAL TEXT OVERRIDE: the only readable words anywhere are exactly [\"EEG\", \"N170\", \"MEG\", \"M170\", \"≈ 170 ms\"]; render each exactly once with correct Vietnamese diacritics; no other text, English, numbers, pseudo-text, or watermark.",
      "visible_text": [
        "EEG",
        "N170",
        "MEG",
        "M170",
        "≈ 170 ms"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_014.png"
    },
    {
      "id": 15,
      "text": "Các vùng khác có thể góp phần tùy theo thứ bạn tưởng mình nhìn thấy. Superior temporal sulcus, hay rãnh thái dương trên, thường được liên hệ với những tín hiệu khuôn mặt có tính thay đổi như hướng nhìn, chuyển động miệng và biểu cảm. Amygdala tham gia đánh giá ý nghĩa cảm xúc và mức độ nổi bật, đặc biệt khi tín hiệu có vẻ đe dọa hoặc đáng chú ý. Nếu mặt tiền một chiếc xe trông “giận dữ”, trải nghiệm ấy không chỉ đến từ hai đèn và một lưới tản nhiệt. Góc nghiêng giống lông mày, độ tương phản và kiến thức xã hội về vẻ giận dữ cùng tác động lên cách bạn cảm nhận.",
      "start": 307.26,
      "end": 338.9,
      "duration": 31.64,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Network-based emotion and motion analysis using an angry-looking car front. Foreground: the recurring white round-headed protagonist in an orange shirt studies headlights angled like eyebrows and a grille shaped like a tense mouth, while expression arrows map these cues onto an angry face template. Midground: a brain diagram highlights superior temporal sulcus for changing facial signals such as gaze, mouth motion, and expression, and highlights the amygdala for emotional salience and possible threat. Background: comparison vehicles vary headlight angle, contrast, tilt, and grille shape, producing neutral, friendly, or angry impressions despite remaining machines. Include at least four narration details: STS involvement, gaze direction, mouth movement, expression, amygdala salience, apparent threat, eyebrow-like angle, contrast, and learned social knowledge. Explanatory device: feature-to-network flowchart with separate arrows from dynamic cues to STS and emotional cues to amygdala. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"STS\", \"Hạch hạnh nhân\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. SCIENTIFIC AND COMPOSITION CORRECTION: Use a mandatory intermediate sequence: car headlights and grille -> abstract face-like eye-mouth configuration -> broad face/social processing icon. Only after that split into two branches: STS receives icons for changing gaze, mouth motion, and inferred social intention; amygdala receives icons for salience, ambiguity, and emotional relevance. No direct arrow from a car to STS or amygdala. Do not portray the amygdala as only fear or anger. FINAL TEXT OVERRIDE: the only readable words anywhere are exactly [\"STS\", \"Hạch hạnh nhân\"]; render each exactly once with correct Vietnamese diacritics; no other text, English, numbers, pseudo-text, or watermark.",
      "visible_text": [
        "STS",
        "Hạch hạnh nhân"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_015.png"
    },
    {
      "id": 16,
      "text": "Công trình của Doris Tsao và cộng sự trên linh trưởng đã làm rõ hơn cách những vùng chuyên đáp ứng với khuôn mặt có thể mã hóa các chiều đặc điểm khuôn mặt theo một hệ thống có cấu trúc. Những kết quả ấy không đồng nghĩa rằng từng “tế bào khuôn mặt” đơn lẻ chứa trọn vẹn hình ảnh của một người, cũng không thể được chuyển thẳng sang mọi trải nghiệm ở người. Chúng gợi ý rằng nhận diện khuôn mặt dựa trên các quần thể nơron xử lý nhiều thuộc tính và quan hệ giữa thuộc tính.",
      "start": 338.9,
      "end": 364.06,
      "duration": 25.16,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Population-coding explainer inspired by Doris Tsao's primate research. Foreground: the recurring white round-headed protagonist in an orange shirt manipulates sliders for eye spacing, face width, head orientation, feature height, and expression while a family of schematic faces changes systematically. Midground: many simplified neuron icons each respond to different dimensions and combine into a population pattern; no single neuron contains a complete portrait. Background: a cautious comparison bridge links primate face patches to human face perception without treating the findings as identical or universally transferable. Include at least four narration details: Doris Tsao and colleagues, primate research, face-responsive regions, structured feature dimensions, population coding, no single complete-face cell, and limits of direct generalization to humans. Explanatory device: multidimensional coordinate diagram feeding a matrix of neurons, followed by a reconstructed face representation. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Doris Tsao\", \"Mã hóa quần thể\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Doris Tsao",
        "Mã hóa quần thể"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_016.png"
    },
    {
      "id": 17,
      "text": "Khi vật vô tri tình cờ khớp một phần quan hệ đó, hệ thống có thể phản ứng dù đầu vào không phải mặt thật.",
      "start": 364.06,
      "end": 369.58,
      "duration": 5.52,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Concise relational-matching demonstration that remains visually information-rich. Foreground: the recurring white round-headed protagonist in an orange shirt aligns two circular bolts and one curved vent on an industrial panel with a transparent triangular face template. Midground: separate feature cards fail to trigger strongly until their spacing and relative positions match the learned two-eyes-over-mouth relationship. Background: a neural population meter rises even though an object-category badge still identifies the input as a machine panel. Include at least four narration details: an inanimate object, partial feature matching, importance of relations between features, activation despite non-face input, preserved object identity, and a face-system response. Explanatory device: before-and-after alignment diagram with measuring lines, relational arrows, and an activation gauge. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_017.png"
    },
    {
      "id": 18,
      "text": "Twist nằm ở đây: bạn không nhìn thấy khuôn mặt vì hệ thị giác yếu kém. Bạn nhìn thấy nó phần nào vì hệ thị giác rất giỏi. Một bộ dò chỉ hoạt động với khuôn mặt hoàn hảo dưới ánh sáng hoàn hảo sẽ thất bại trong đời thực, nơi mặt người quay nghiêng, bị tóc che, chìm trong bóng tối, xuất hiện thoáng qua hoặc chỉ chiếm vài điểm ảnh. Để nhận ra người thật trong điều kiện khó, hệ thống phải tổng quát hóa. Chính khả năng tổng quát hóa hữu ích ấy đôi lúc vượt quá mục tiêu và biến ba hình đơn giản thành một nhân vật có cảm xúc.",
      "start": 369.58,
      "end": 399.06,
      "duration": 29.48,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Educational twist scene showing that pareidolia emerges from useful visual competence. Foreground: the recurring white round-headed protagonist in an orange shirt stands before a detector that successfully recognizes real faces when they are turned sideways, partly covered by hair, shadowed, fleeting, or reduced to a few pixels. Midground: the same generalizing detector occasionally classifies three simple marks as an emotional character, producing a harmless overshoot indicator rather than a malfunction warning. Background: a comparison detector requiring perfect frontal lighting misses most real-world faces, while the flexible detector succeeds across difficult conditions. Include at least four narration details: vision is highly capable, perfect-only detection would fail, faces can be rotated, occluded, dark, brief, or pixelated, generalization is necessary, and useful generalization can overshoot. Explanatory device: side-by-side performance chart for rigid versus flexible detectors with arrows from varied inputs to outcomes. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_018.png"
    },
    {
      "id": 19,
      "text": "Thế giới hiện đại đang cung cấp cho hệ thống này nhiều bề mặt mới. Camera điện thoại tăng tương phản, làm nét, ghép nhiều khung hình và đôi lúc biến nhiễu thành đường nét có vẻ có chủ ý. Bộ lọc khuôn mặt có thể đặt tai, mắt kính hoặc lớp trang điểm lên một cái gối vì thuật toán cũng đang tìm một cấu hình gần giống hai mắt và một miệng. Ảnh thumbnail thu nhỏ làm mất chi tiết nhưng giữ bố cục lớn, chính xác là điều kiện thuận lợi cho pareidolia.",
      "start": 399.06,
      "end": 423.74,
      "duration": 24.68,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Modern digital pareidolia workstation in a detailed educational stick-figure style. Foreground: the recurring white round-headed protagonist in an orange shirt aims a phone camera at a patterned pillow; a face filter mistakenly adds ears, glasses, and makeup to two buttons and a seam. Midground: a processing pipeline shows contrast enhancement, sharpening, multi-frame merging, and noise becoming deliberate-looking edges, followed by a thumbnail that loses details but preserves a face-like global layout. Background: a wall of compressed social-media thumbnails and low-light camera frames offers many new ambiguous surfaces for human and algorithmic detectors. Include at least four narration details: phone-camera contrast enhancement, sharpening, frame merging, noise-to-edge conversion, filter misplacement on a pillow, and detail loss in thumbnails. Explanatory device: sequential image-processing flowchart with magnified before-and-after crops and detector bounding boxes. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Bộ lọc AI\", \"Ảnh thu nhỏ\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Bộ lọc AI",
        "Ảnh thu nhỏ"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_019.png"
    },
    {
      "id": 20,
      "text": "Emoji lại chứng minh mức tối giản của tín hiệu xã hội: vài dấu chấm và một đường cong vẫn đủ để bạn đọc ra vui vẻ, mỉa mai hoặc thất vọng.",
      "start": 423.74,
      "end": 432.62,
      "duration": 8.88,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Minimal-signal social decoding board that is dense in comparisons rather than empty space. Foreground: the recurring white round-headed protagonist in an orange shirt rearranges two dots and one curved line into several emoji-like configurations that instantly suggest joy, sarcasm, disappointment, and neutrality. Midground: nearby stick figures respond differently to each tiny configuration despite the same limited number of marks. Background: a feature economy chart reduces a detailed human face step by step to eyes and a mouth while preserving recognizable emotional information. Include at least four narration details: very few marks, two eye-like dots, one mouth curve, happiness, sarcasm, disappointment, and rapid social interpretation. Explanatory device: transformation sequence and expression matrix connecting line curvature and eye placement to perceived emotion. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_020.png"
    },
    {
      "id": 21,
      "text": "Bạn có thể kiểm tra trải nghiệm bằng vài thay đổi rất đơn giản. Hãy đổi góc nhìn: bước sang bên hoặc xoay vật thể. Nếu gương mặt phụ thuộc vào một sự thẳng hàng tình cờ, nó sẽ tan rã. Hãy thay ánh sáng: bật thêm đèn, kéo rèm hoặc quan sát vào thời điểm khác. Nếu đôi mắt biến mất cùng bóng đổ, nguyên nhân đã rõ hơn. Hãy thay khoảng cách: tiến gần để lấy lại chi tiết, hoặc lùi xa để xem cấu hình tổng quát mạnh đến mức nào. Bạn cũng có thể che lần lượt một trong hai điểm giống mắt hoặc đường giống miệng. Chỉ cần mất một mốc, cảm giác về mặt thường suy yếu đáng kể.",
      "start": 432.62,
      "end": 464.34,
      "duration": 31.72,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Hands-on four-test laboratory featuring the recurring white round-headed protagonist in an orange shirt investigating one face-like household object. Foreground: the protagonist performs four actions in sequence—steps sideways to change angle, rotates the object, switches on a lamp and opens a curtain, then moves closer and farther away. Midground: a final test covers one eye-like point, the second eye-like point, and the mouth-like line in turn, visibly weakening the face impression. Background: comparison snapshots show accidental alignment breaking, shadow-eyes disappearing, fine detail returning at close range, and broad configuration strengthening at distance. Include at least four narration details: angle change, object rotation, lighting change, curtain or added lamp, distance variation, recovering detail, and covering individual facial landmarks. Explanatory device: four-panel test protocol with arrows, before-and-after states, and a face-strength meter. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Đổi góc\", \"Đổi sáng\", \"Đổi khoảng cách\", \"Che điểm\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. SCIENTIFIC AND COMPOSITION CORRECTION: Render the four Vietnamese labels exactly. Show the same face-like cabinet/outlet tested by changing viewing angle, changing light level, changing distance, and covering one eye-like point at a time; use a four-step comparison with a clear before-and-after conclusion. FINAL TEXT OVERRIDE: the only readable words anywhere are exactly [\"Đổi góc\", \"Đổi độ sáng\", \"Thay đổi khoảng cách\", \"Che từng điểm\"]; render each exactly once with correct Vietnamese diacritics; no other text, English, numbers, pseudo-text, or watermark. FINAL REGENERATION OVERRIDE: Remove all numbered badges, digits, step numbers, punctuation marks, and extra symbols. Use four horizontal rows distinguished by color and the four exact permitted Vietnamese labels only. Preserve the angle, light, distance, and cover-one-point tests.",
      "visible_text": [
        "Đổi góc",
        "Đổi độ sáng",
        "Thay đổi khoảng cách",
        "Che từng điểm"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_021.png"
    },
    {
      "id": 22,
      "text": "Một cách kiểm tra khác là chụp ảnh rồi xoay ảnh ngược. Khả năng nhận diện khuôn mặt thật thường bị ảnh hưởng khi khuôn mặt đảo chiều vì quan hệ quen thuộc giữa các đặc điểm bị phá vỡ. Gương mặt pareidolia cũng có thể mất sức thuyết phục khi hướng bị đổi, dù kết quả tùy hình. Bạn cũng có thể hỏi một người khác họ nhìn thấy gì trước khi gợi ý. Nếu phải chỉ từng “con mắt” và “cái miệng” rất lâu, bằng chứng trong hình có lẽ khá yếu và kỳ vọng của bạn đang làm nhiều việc hơn.",
      "start": 464.34,
      "end": 491.02,
      "duration": 26.68,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Reality-testing experiment in an educational stick-figure style. Foreground: the recurring white round-headed protagonist in an orange shirt photographs a face-like object, rotates the image upside down, and watches the facial impression weaken. Midground: a second stick figure independently reports what they see before receiving any hint, while the protagonist keeps eye and mouth annotations hidden. Background: an inversion comparison shows familiar feature relations disrupted in a real face and a pareidolia example, and an evidence gauge drops when recognition requires prolonged prompting. Include at least four narration details: photograph capture, image inversion, disruption of familiar feature relations, variable effect across images, asking another person without suggestion, and weak evidence when landmarks need extensive pointing. Explanatory device: upright-versus-inverted split-screen, blinded-observer workflow, and expectation-versus-image evidence balance. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Xoay ngược\", \"Hỏi trước khi gợi ý\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. SCIENTIFIC AND COMPOSITION CORRECTION: Render the two Vietnamese labels exactly. Show a photographed face-like object, the image rotated 180 degrees, then a separate memory sequence comparing recall before and after a visual hint. Do not use any other words. FINAL TEXT OVERRIDE: the only readable words anywhere are exactly [\"Lật ảnh\", \"Nhớ lại sau gợi ý\"]; render each exactly once with correct Vietnamese diacritics; no other text, English, numbers, pseudo-text, or watermark.",
      "visible_text": [
        "Lật ảnh",
        "Nhớ lại sau gợi ý"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_022.png"
    },
    {
      "id": 23,
      "text": "Pareidolia thông thường thường gắn với một kích thích có thật: vệt bẩn, bóng cây, hoa văn, tiếng nhiễu. Bạn vẫn có thể kiểm tra, đổi góc và nhận ra cách hiểu khác. Hallucination, hay ảo giác theo nghĩa lâm sàng, là trải nghiệm cảm giác sống động khi không có kích thích bên ngoài tương ứng, hoặc trải nghiệm dai dẳng khó điều chỉnh bằng kiểm tra thực tế. Ranh giới trong đời sống không phải lúc nào cũng đơn giản, và một hiện tượng đơn lẻ không đủ để tự chẩn đoán.",
      "start": 491.02,
      "end": 517.62,
      "duration": 26.6,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Careful clinical-boundary comparison with calm, non-alarmist educational visuals. Foreground: the recurring white round-headed protagonist in an orange shirt examines a real stain that resembles a face, changes viewing angle, and recognizes the stain underneath. Midground: a separate panel depicts a vivid percept with no corresponding external stimulus and another that persists despite repeated reality checks, without sensational imagery. Background: a continuum includes ambiguous patterns, testable pareidolia, persistent unsupported perception, and a warning against self-diagnosis from one isolated event. Include at least four narration details: real external stimulus such as stain or shadow, ability to inspect and reinterpret it, hallucination without matching external input, vivid or persistent experience, difficulty updating through reality checks, and a non-simple everyday boundary. Explanatory device: side-by-side stimulus-present versus stimulus-absent diagram plus a graded continuum rather than a hard binary wall. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Có kích thích\", \"Không có kích thích\", \"Không tự chẩn đoán\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Có kích thích",
        "Không có kích thích",
        "Không tự chẩn đoán"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_023.png"
    },
    {
      "id": 24,
      "text": "Nếu bạn chỉ thỉnh thoảng thấy mặt trong ổ cắm, mây hoặc bọt cà phê và biết đó là đồ vật, điều này thường hoàn toàn bình thường. Nhưng nếu các khuôn mặt hoặc hình người xuất hiện thường xuyên khi không có mẫu bên ngoài, gây sợ hãi, ra lệnh, làm gián đoạn giấc ngủ hay sinh hoạt, đi kèm lú lẫn, suy giảm trí nhớ, thay đổi hành vi, sốt, chấn thương, dùng chất hoặc thay đổi thuốc, việc trao đổi với bác sĩ hay chuyên gia sức khỏe tâm thần là phù hợp. Nếu sự thay đổi xảy ra đột ngột và nghiêm trọng, cần tìm hỗ trợ y tế sớm thay vì chỉ xem đó là pareidolia.",
      "start": 517.62,
      "end": 550.08,
      "duration": 32.46,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Balanced normal-versus-seek-help decision guide in a respectful educational stick-figure style. Foreground: the recurring white round-headed protagonist in an orange shirt calmly recognizes occasional faces in an outlet, cloud, and coffee foam while maintaining object labels and normal daily activity. Midground: a contrasting pathway shows frequent unsupported figures causing fear, commands, sleep disruption, or impaired routine, followed by a supportive conversation with a doctor or mental-health professional. Background: additional medical-context icons show confusion, memory decline, behavior change, fever, injury, substance use, and medication change; a rapid-onset severe branch leads to early medical support. Include at least four narration details: occasional object-based faces are usually normal, preserved insight, frequent perception without external patterns, distress or commanding content, disrupted sleep or function, accompanying medical changes, and sudden severe onset. Explanatory device: two-path decision tree with green monitoring icons, amber consultation icons, and a direct urgent-support arrow. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. The only readable words anywhere are exactly [\"Thường bình thường\", \"Nên trao đổi chuyên gia\", \"Đột ngột, nghiêm trọng\"]. Render each phrase exactly once with correct Vietnamese diacritics; no other letters, labels, numbers, English, pseudo-text, or watermark. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. SCIENTIFIC AND COMPOSITION CORRECTION: Build three unmistakable lanes: green normal pareidolia with insight preserved; amber repeated distress or daily impairment leading to a scheduled conversation with a clinician; red sudden severe confusion, fever, head injury, medication/substance change, or inability to stay safe leading to an emergency medical endpoint. The red lane must end at a hospital/emergency icon, not the same routine consultation. Render only the three exact Vietnamese labels. FINAL TEXT OVERRIDE: the only readable words anywhere are exactly [\"Thường là bình thường\", \"Nên trao đổi với chuyên gia\", \"Tìm trợ giúp y tế ngay\"]; render each exactly once with correct Vietnamese diacritics; no other text, English, numbers, pseudo-text, or watermark. FINAL REGENERATION OVERRIDE: Remove every exclamation mark, question mark, punctuation-only warning badge, and any characters outside the three exact permitted phrases. Use icon shapes such as solid red circles and hospital symbols without punctuation. Preserve three distinct green, amber, and red lanes and the separate emergency hospital endpoint. ABSOLUTE SYMBOL OVERRIDE: remove every question mark, exclamation mark, ellipsis, checkmark, punctuation badge, speech-bubble dots, and writing-like mark. Emotions must be shown only by eyebrows, mouth curves, posture, and solid-color icons. Keep each of the three permitted Vietnamese phrases exactly once and no other characters. LAST SYMBOL-FREE OVERRIDE: leave completely empty blank space around and above every character head. No floating strokes, spirals, swirls, scribbles, punctuation, bubbles, rays, stars, badges, marks, or symbols near any person. Express distress only through body posture and eyebrows; express medical conditions only with recognizable objects such as bed, thermometer without scale marks, bandage, medicine bottle without label, ambulance, and hospital building. Keep exactly the three permitted phrases and no other writing-like shapes.",
      "visible_text": [
        "Thường là bình thường",
        "Nên trao đổi với chuyên gia",
        "Tìm trợ giúp y tế ngay"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_024.png"
    },
    {
      "id": 25,
      "text": "Lần tới khi cụm đèn xe nhìn bạn với vẻ khó chịu, bạn có thể dừng lại một nhịp. Hai vùng sáng tạo thành mắt. Lưới tản nhiệt tạo thành miệng. Tầm nhìn, bóng đổ, ký ức về hàng nghìn khuôn mặt và ngưỡng phát hiện thấp đã gặp nhau trong khoảnh khắc. Cảm giác ấy vừa thật với trải nghiệm của bạn, vừa không phải bằng chứng rằng vật thể có tâm trí. Đó là một dự đoán hữu ích được đưa ra quá hào phóng.",
      "start": 550.08,
      "end": 571.78,
      "duration": 21.7,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Synthesis scene returning to an angry-looking car. Foreground: the recurring white round-headed protagonist in an orange shirt pauses beside the car and analytically points to two bright headlights as eyes and the grille as a mouth, while a translucent angry-face overlay appears but does not replace the vehicle. Midground: arrows bring together viewing angle, cast shadows, memories of thousands of faces, and a low detection threshold inside a brain-shaped inference hub. Background: the hub outputs a socially meaningful prediction while a separate object-reality track confirms that the car has no mind or emotion. Include at least four narration details: annoyed-looking lights, eye-like bright regions, mouth-like grille, role of viewpoint, role of shadow, accumulated face memory, low threshold, and real subjective feeling without evidence of mentality. Explanatory device: converging causal diagram ending in an over-generous prediction meter. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_025.png"
    },
    {
      "id": 26,
      "text": "Rồi bạn trở lại bàn làm việc. Chiếc ổ cắm dưới chân tường vẫn có hai lỗ tròn và một khe nhỏ. Có thể nó vẫn trông hơi ngạc nhiên, nhất là khi bạn nhìn từ khóe mắt. Nhưng nó vẫn chỉ là ổ cắm. Trong phần nhỏ của một giây, não bạn đã thử một giả thuyết xã hội cực nhanh: có một khuôn mặt ở đó không? Câu trả lời là không, nhưng khả năng đặt câu hỏi ấy nhanh đến vậy chính là một phần lý do bạn nhận ra khuôn mặt thật giữa một thế giới đầy nhiễu.",
      "start": 571.78,
      "end": 595.16,
      "duration": 23.38,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Circular closing composition matching the opening workspace. Foreground: the recurring white round-headed protagonist in an orange shirt returns to the desk, glances from peripheral vision toward the same wall outlet, then kneels to inspect its two round holes and one narrow slot as ordinary hardware. Midground: a rapid brain hypothesis panel asks whether the three marks form a social face template, then rejects the hypothesis after direct inspection while preserving the speed of detection. Background: faint callbacks to the real human faces successfully detected in crowds, darkness, occlusion, and visual noise connect this harmless outlet false alarm to useful everyday recognition. Include at least four narration details: return to the desk, unchanged outlet geometry, stronger surprise impression in peripheral vision, correct object conclusion, sub-second social hypothesis, negative answer after checking, and rapid real-face recognition in noise. Explanatory device: looped timeline from opening outlet to scientific explanation to final outlet, with a fast hypothesis-check-result flowchart. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. No readable words, letters, labels, numbers, pseudo-text, or watermark anywhere; use icons and diagrams only. Use only clean uniform fill regions from the fixed palette; every background, object, screen, face-like pattern, brain, shirt, and diagram must be one solid color with hard boundaries. No glow, haze, transparency, grain, noise, speckles, highlights, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. SCIENTIFIC AND COMPOSITION CORRECTION: Create a calm final return to the original desk. The unchanged wall outlet has two holes and one slot; the orange-shirt protagonist observes it from a safe distance. A simple icon-only flow reads: ambiguous outlet pattern -> quick social face hypothesis in a thought bubble -> visual check -> hypothesis bubble opens and dissolves -> clear outlet object, then a forward arrow toward detecting a real human face in visual noise. No tools, no screwdriver, no repair, no touching the outlet, no electricity symbol, no question mark, no checkmark, no X, no letters, numbers, pseudo-text, or labels. Keep one protagonist only and make the ending direction unambiguous. FINAL TEXT OVERRIDE: no readable words, letters, labels, numbers, punctuation-like marks, pseudo-text, or watermark anywhere. FINAL REGENERATION OVERRIDE: Make a left-to-right ending with no loops and no extra crowd vignettes. Show exactly one calm orange-shirt protagonist seated at the original desk, safely looking at the unchanged outlet. Then a single straight sequence of five icon panels: outlet pattern, temporary face hypothesis cloud, eye checking, cloud dissolving back into the outlet, real human face in visual noise. End the final arrow at the real face; no arrow returns. Calm eyebrows and closed relaxed line mouth. No tools, touching, electricity, punctuation, text, checkmarks, X marks, or extra people outside the final real-face icon.",
      "visible_text": [],
      "output_path": "/data/video-pipeline/StickerMan/project/005-Vi-Sao-Ban-De-Nhin-Thay-Khuon-Mat-Trong-Nhung-Vat-Vo-Tri/images/scenes/scene_026.png"
    }
  ]
}
