{
  "version": 1,
  "language": "vi",
  "items": [
    {
      "id": 1,
      "text": "Bạn đang ngồi một mình trong căn hộ vào buổi tối. Màn hình laptop sáng trước mặt, tai nghe phát một đoạn nhạc quen thuộc, ngoài cửa sổ chỉ có vài ô cửa còn bật đèn. Không có tiếng bước chân, không có thông báo từ camera, cũng không có lý do rõ ràng để lo lắng. Thế rồi, giữa một khoảnh khắc hoàn toàn bình thường, bạn bỗng thấy gáy mình căng lên. Một ý nghĩ xuất hiện rất nhanh: có ai đó đang nhìn mình.",
      "start": 0.0,
      "end": 21.86,
      "duration": 21.86,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Information-rich educational stick-figure scene in a dark apartment. Foreground: the recurring white round-headed protagonist in an orange shirt sits at a laptop with a single flat pale-yellow screen rectangle, wears headphones, and suddenly stiffens as short tension lines rise from the neck. Midground: a locked door, a motionless curtain, and a silent security-camera panel establish that no warning is present. Background: a flat navy night window shows only a few illuminated apartments and a solid flat blue silhouette behind one window pane. Use a thought bubble containing a single staring eye and a dashed arrow from the solid silhouette toward the protagonist. End with the protagonist beginning to turn, continuing directly into scene 2. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Có ai đang nhìn mình?\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. The security camera panel contains only four simple picture cells and one solid green status dot; absolutely no letters, words, numbers, status text, English text, or pseudo-text on the panel. The only readable words anywhere in the entire image are exactly: “Có ai đang nhìn mình?”",
      "visible_text": [
        "Có ai đang nhìn mình?"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_001.png"
    },
    {
      "id": 2,
      "text": "Bạn quay lại. Căn phòng vẫn trống. Tấm rèm không động đậy. Cánh cửa vẫn khóa. Có thể bạn tự cười, cho rằng não vừa tưởng tượng quá mức. Nhưng cảm giác ấy từng rõ đến mức gần như một giác quan riêng, như thể cơ thể đã nhận được một tín hiệu mà ý thức chưa kịp hiểu.",
      "start": 21.86,
      "end": 37.0,
      "duration": 15.14,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Continue inside the same dark apartment with the same white round-headed protagonist in an orange shirt completing the turn. Foreground: the protagonist looks over one shoulder, half nervous and half embarrassed, while a small body-signal diagram marks tension at the neck before conscious reasoning. Midground: the room is visibly empty, the curtain remains still, and the door remains locked. Background: the laptop glow and unchanged night window preserve the exact geography from scene 1. Add a split-screen inset contrasting 'strong bodily signal' with 'no visible observer,' linked by a question-mark arrow. Pull the inset outward into the multi-signal explanation of scene 3. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Cảm giác rất thật\", \"Phòng vẫn trống\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"CẢM GIÁC RẤT THẬT\", \"PHÒNG VẪN TRỐNG\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "CẢM GIÁC RẤT THẬT",
        "PHÒNG VẪN TRỐNG"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_002.png"
    },
    {
      "id": 3,
      "text": "Câu hỏi quen thuộc là: liệu con người có thật sự cảm nhận được ánh mắt từ phía sau hay không? Nhưng câu trả lời thật sự còn kỳ lạ hơn. Não không cần có một năng lực bí ẩn để khiến bạn chắc chắn rằng mình đang bị quan sát. Nó chỉ cần vài dấu hiệu mơ hồ, một hệ thống cảnh báo nhạy quá mức, và một cỗ máy dự đoán luôn cố tìm ra tác nhân đứng sau mọi biến động trong môi trường.",
      "start": 37.0,
      "end": 57.66,
      "duration": 20.66,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Transform the apartment into an educational cutaway while keeping the orange-shirt protagonist centered. Foreground: the protagonist asks whether a person can sense a gaze from behind as three translucent inputs approach the head: a vague shadow, a tiny sound wave, and peripheral motion. Midground: an oversized brain-shaped warning console combines ambiguous clues, a sensitive alarm bell, and a prediction engine that searches for an agent. Background: the dark apartment remains faintly visible to show that no supernatural observer is required. Use a flow diagram reading ambiguous cues to over-sensitive warning to predicted watcher, with a crossed-out magic-eye icon. Carry the incoming cue lines into the eye-gaze laboratory of scene 4. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Không cần giác quan bí ẩn\", \"Dấu hiệu → Cảnh báo → Dự đoán\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Không cần giác quan bí ẩn",
        "Dấu hiệu → Cảnh báo → Dự đoán"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_003.png"
    },
    {
      "id": 4,
      "text": "Cảm giác bị nhìn không bắt đầu từ một “máy dò ánh mắt” duy nhất. Nó xuất hiện từ nhiều hệ thống phối hợp với nhau. Não liên tục đọc hướng đầu, vị trí khuôn mặt, độ tương phản quanh mắt, chuyển động trong tầm nhìn ngoại vi, âm thanh nhỏ phía sau và cả ký ức về nơi bạn đang đứng. Phần lớn quá trình này diễn ra trước khi bạn kịp suy nghĩ thành lời.",
      "start": 57.66,
      "end": 75.92,
      "duration": 18.26,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Enter a clean eye-gaze laboratory, with the same orange-shirt protagonist seated at a scanning station. Foreground: five labeled icon streams enter a transparent head model: head direction, face position, contrast around the eyes, peripheral movement, and a small sound from behind. Midground: a memory card showing the dark apartment joins those live signals before they reach awareness. Background: researchers monitor a simplified brain display while the earlier apartment is shown as a small continuity thumbnail. Use a layered processing diagram in which the five inputs converge below a speech bubble, emphasizing that integration occurs before verbal thought. Finish by enlarging the face-and-eye input for scene 5. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Nhiều hệ thống phối hợp\", \"Xử lý trước ý thức\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Nhiều hệ thống phối hợp",
        "Xử lý trước ý thức"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_004.png"
    },
    {
      "id": 5,
      "text": "Ánh mắt đặc biệt quan trọng vì nó tiết lộ sự chú ý của người khác. Nếu một người nhìn thẳng vào bạn, họ có thể muốn giao tiếp, đánh giá, hợp tác, cạnh tranh hoặc chuẩn bị hành động. Trong lịch sử tiến hóa, nhận ra hướng nhìn nhanh hơn một phần giây có thể tạo ra khác biệt đáng kể. Vì vậy, hệ thần kinh có xu hướng ưu tiên khuôn mặt và đôi mắt, nhất là ánh mắt hướng trực tiếp về phía bạn.",
      "start": 75.92,
      "end": 98.74,
      "duration": 22.82,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Zoom from the laboratory display into an information-rich social-gaze tableau. Foreground: the orange-shirt protagonist faces five stick figures whose direct gazes imply communication, evaluation, cooperation, competition, and imminent action through distinct hand poses and symbols. Midground: a large pair of eyes points directly toward the protagonist while a small reaction clock shows a fraction-of-a-second advantage. Background: an evolutionary strip moves from an ancestral camp to a modern crowd, showing why rapid gaze detection matters across contexts. Use a radial diagram connecting direct gaze to the five possible intentions, with the eye route highlighted above other facial features. Transition by turning the radial diagram into a controlled visual-search test in scene 6. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Ánh mắt tiết lộ sự chú ý\", \"Ưu tiên ánh mắt trực diện\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Ánh mắt tiết lộ sự chú ý",
        "Ưu tiên ánh mắt trực diện"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_005.png"
    },
    {
      "id": 6,
      "text": "Hiện tượng này thường được gọi là lợi thế của ánh mắt trực diện, hay direct-gaze advantage. Trong nhiều thí nghiệm thị giác, người tham gia phát hiện khuôn mặt đang nhìn thẳng nhanh hơn hoặc xử lý nó khác với khuôn mặt nhìn lệch. Ánh mắt trực diện cũng dễ thu hút chú ý và làm tăng cảm giác rằng thông tin đang liên quan đến bản thân. Hiệu ứng không giống nhau trong mọi điều kiện, nhưng kết luận tổng quát khá vững: hướng nhìn là một tín hiệu xã hội có mức ưu tiên cao.",
      "start": 98.74,
      "end": 123.64,
      "duration": 24.9,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Continue in the eye-gaze lab with the orange-shirt protagonist performing a visual-search experiment. Foreground: a grid of simplified faces contains one face looking directly outward among several faces looking sideways, and the protagonist presses a response button as the direct face lights up first. Midground: a comparison panel shows direct gaze capturing attention, receiving different processing, and increasing self-relevance. Background: researchers record several conditions, including one where the effect is weaker, preventing an absolute claim. Use a race-track diagram with a direct-gaze lane reaching attention before an averted-gaze lane, plus a modest variability marker. The highlighted face morphs into a brain scan focused on the temporal region for scene 7. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Lợi thế ánh mắt trực diện\", \"Tín hiệu xã hội ưu tiên cao\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Lợi thế ánh mắt trực diện",
        "Tín hiệu xã hội ưu tiên cao"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_006.png"
    },
    {
      "id": 7,
      "text": "Một vùng não quan trọng trong quá trình này là rãnh thái dương trên, thường được gọi bằng tên tiếng Anh là superior temporal sulcus, hay STS. Vùng này tham gia xử lý nhiều tín hiệu xã hội động, bao gồm hướng nhìn, chuyển động đầu, biểu cảm và ý định được suy ra từ hành động. Khi mắt người khác đổi hướng, STS giúp não theo dõi xem sự chú ý của họ đang hướng vào đâu. Nó góp phần biến một chuyển động rất nhỏ của đồng tử và mí mắt thành thông tin xã hội có ý nghĩa.",
      "start": 123.64,
      "end": 151.28,
      "duration": 27.64,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Reveal a large side-view brain diagram beside the same orange-shirt protagonist in the laboratory. Foreground: a glowing groove in the upper temporal region is identified as STS while animated eye direction, head movement, facial expression, and inferred action intention feed into it. Midground: STS converts a tiny pupil-and-eyelid shift into an arrow showing where another person is attending. Background: a social scene demonstrates one figure turning their eyes toward an object while the protagonist follows that attention. Use a four-input-to-one-social-meaning diagram with moving arrows and a magnified eye inset. Continue along the outgoing attention arrow into the gaze-cueing experiment of scene 8. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"STS\", \"Hướng nhìn\", \"Chuyển động đầu\", \"Ý định\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "STS",
        "Hướng nhìn",
        "Chuyển động đầu",
        "Ý định"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_007.png"
    },
    {
      "id": 8,
      "text": "Các nghiên cứu về gaze cueing, tức hiệu ứng gợi hướng bằng ánh mắt, cho thấy sức mạnh của tín hiệu này. Trong một thí nghiệm điển hình, bạn nhìn thấy một khuôn mặt ở giữa màn hình. Đôi mắt của khuôn mặt nhìn sang trái hoặc phải. Sau đó, một mục tiêu xuất hiện ở một bên. Bạn thường phản ứng nhanh hơn khi mục tiêu xuất hiện ở hướng mà đôi mắt vừa nhìn, ngay cả khi hướng nhìn không dự đoán chính xác vị trí mục tiêu.",
      "start": 151.28,
      "end": 171.14,
      "duration": 19.86,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Continue the laboratory sequence with a classic gaze-cueing trial. Foreground: the orange-shirt protagonist watches a central face whose eyes shift left, then reacts to a target appearing on one side. Midground: split the screen into congruent and incongruent trials, showing a faster button press when the target follows the gaze and a slower response when it appears opposite. Background: a randomized trial board makes clear that eye direction does not reliably predict target location. Use a step-by-step timeline diagram: central face, eye shift, target onset, reaction, with arrows and compact reaction bars. Preserve the eye-arrow motif and pass it to the research context in scene 9. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Gợi hướng bằng ánh mắt\", \"Nhanh hơn\", \"Chậm hơn\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Gợi hướng bằng ánh mắt",
        "Nhanh hơn",
        "Chậm hơn"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_008.png"
    },
    {
      "id": 9,
      "text": "Andrew Bayliss và Steven Tipper đã công bố nhiều nghiên cứu về cách ánh mắt tự động định hướng sự chú ý và còn ảnh hưởng đến việc đánh giá người khác. Kết quả thuộc dòng nghiên cứu này gợi ý rằng bạn không chỉ nhìn thấy ánh mắt; hệ chú ý còn dùng ánh mắt như một mũi tên xã hội. Đôi mắt của người khác có thể kéo sự chú ý của bạn về một hướng trước khi ý thức quyết định có nên làm vậy hay không.",
      "start": 171.14,
      "end": 182.9,
      "duration": 11.76,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Expand the gaze-cueing setup into a research-board scene while retaining the orange-shirt protagonist and central face. Foreground: the protagonist's attention spotlight is pulled sideways by another person's eyes before a conscious decision bubble appears. Midground: stylized researcher portraits and paper cards represent Andrew Bayliss and Steven Tipper, alongside icons for automatic orienting and person evaluation. Background: repeated trials show eyes functioning like social arrows toward objects and people. Use an explanatory overlay that literally converts the pupils into a directional arrow entering an attention map, with a small 'before awareness' sequence. Darken the arrow's destination into an emotionally threatening face, setting up the amygdala in scene 10. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Mũi tên xã hội\", \"Chú ý tự động\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Mũi tên xã hội",
        "Chú ý tự động"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_009.png"
    },
    {
      "id": 10,
      "text": "Một cấu trúc khác thường được nhắc đến là hạch hạnh nhân, hay amygdala. Amygdala không đơn giản là “trung tâm sợ hãi”. Nó tham gia đánh giá mức độ nổi bật và ý nghĩa sinh học của tín hiệu, đặc biệt khi tín hiệu liên quan đến cảm xúc, đe dọa hoặc tương tác xã hội. Ánh mắt trực diện từ một khuôn mặt giận dữ, chẳng hạn, có thể mang ý nghĩa khác hẳn ánh mắt nhìn lệch từ cùng khuôn mặt đó.",
      "start": 182.9,
      "end": 205.02,
      "duration": 22.12,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Shift from the eye-arrow to a brain salience diagram featuring the same orange-shirt protagonist. Foreground: two versions of an angry face appear, one staring directly at the protagonist and one looking away, producing visibly different alert responses. Midground: a glowing amygdala evaluates emotional meaning, biological importance, threat, and social relevance rather than acting as a simple fear switch. Background: neutral, friendly, and threatening social signals are sorted by importance on a laboratory display. Use a comparison matrix crossing facial emotion with gaze direction, plus an arrow from direct angry gaze to the strongest salience marker. Zoom toward the eye-sampling path for scene 11. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Amygdala\", \"Không chỉ là “trung tâm sợ hãi”\", \"Độ nổi bật và ý nghĩa\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"HẠCH HẠNH NHÂN\", \"KHÔNG CHỈ LÀ TRUNG TÂM SỢ HÃI\", \"ĐỘ NỔI BẬT VÀ Ý NGHĨA\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "HẠCH HẠNH NHÂN",
        "KHÔNG CHỈ LÀ TRUNG TÂM SỢ HÃI",
        "ĐỘ NỔI BẬT VÀ Ý NGHĨA"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_010.png"
    },
    {
      "id": 11,
      "text": "Nhà thần kinh học Ralph Adolphs và các cộng sự đã nghiên cứu sâu vai trò của amygdala trong nhận diện cảm xúc và thị giác xã hội. Một số công trình nổi tiếng cho thấy tổn thương amygdala có thể làm giảm khả năng nhận ra nỗi sợ trên khuôn mặt, một phần vì người bệnh không tự nhiên hướng sự chú ý đến vùng mắt. Khi được hướng dẫn nhìn vào mắt, khả năng nhận diện có thể cải thiện. Điều này cho thấy đôi mắt không chỉ chứa thông tin; não còn phải ưu tiên lấy mẫu đúng nơi để sử dụng thông tin ấy.",
      "start": 205.02,
      "end": 229.62,
      "duration": 24.6,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Continue at the brain-and-face workstation with the orange-shirt protagonist observing an eye-tracking demonstration. Foreground: a scan path initially avoids the eyes of a fearful face, reducing emotion recognition, then a guiding arrow directs fixation to the eyes and recognition improves. Midground: a simplified amygdala lesion comparison shows weaker spontaneous sampling of the eye region without implying that information is absent. Background: a research panel represents Ralph Adolphs and colleagues studying emotion recognition and social vision. Use a before-and-after split screen with gaze dots, an eye-region target box, and an accuracy gauge. Let the target box expand into the wide visual field examined in scene 12. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Thông tin nằm ở đôi mắt\", \"Não phải nhìn đúng chỗ\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"THÔNG TIN NẰM Ở ĐÔI MẮT\", \"NÃO PHẢI NHÌN ĐÚNG CHỖ\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "THÔNG TIN NẰM Ở ĐÔI MẮT",
        "NÃO PHẢI NHÌN ĐÚNG CHỖ"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_011.png"
    },
    {
      "id": 12,
      "text": "Tầm nhìn ngoại vi làm cảm giác này mạnh hơn. Ở rìa trường nhìn, bạn không thấy chi tiết tốt như ở trung tâm. Bạn khó phân biệt chính xác đôi mắt, nét mặt hay vật thể nhỏ. Nhưng tầm nhìn ngoại vi lại khá nhạy với chuyển động, thay đổi độ sáng và sự xuất hiện đột ngột. Nó được thiết kế tốt cho câu hỏi “có gì vừa thay đổi?” hơn là câu hỏi “đó chính xác là ai?”.",
      "start": 229.62,
      "end": 248.84,
      "duration": 19.22,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Place the orange-shirt protagonist inside a wide visual-field demonstration. Foreground: the central viewing cone renders a face, eyes, and small objects sharply, while the peripheral edges blur those details. Midground: movement streaks, a sudden brightness change, and a newly appearing silhouette remain conspicuous at the edges. Background: the laboratory wall curves into a panoramic version of the original apartment, reconnecting the science to the opening setting. Use a visual-field diagram contrasting central identification with peripheral change detection and arrows asking 'what changed?' versus 'who exactly?'. A peripheral flicker at the apartment edge becomes the ambiguous objects in scene 13. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Trung tâm: chi tiết\", \"Ngoại vi: thay đổi\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"TRUNG TÂM: CHI TIẾT\", \"NGOẠI VI: THAY ĐỔI\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device. Render each allowed phrase exactly once. “TRUNG TÂM: CHI TIẾT” appears once on the central panel; “NGOẠI VI: THAY ĐỔI” appears once total across the entire image. Do not duplicate either phrase.",
      "visible_text": [
        "TRUNG TÂM: CHI TIẾT",
        "NGOẠI VI: THAY ĐỔI"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_012.png"
    },
    {
      "id": 13,
      "text": "Vì thế, một bóng người phản chiếu trên màn hình tối, chuyển động của rèm, chiếc áo khoác treo trên ghế hoặc người đi ngang ngoài cửa kính có thể tạo ra tín hiệu ban đầu. Não phát hiện chuyển động trước, rồi cố giải thích nguyên nhân. Nếu bối cảnh đã khiến bạn cảnh giác, lời giải thích “có người” thường được ưu tiên hơn “ánh đèn xe vừa quét qua”.",
      "start": 248.84,
      "end": 266.3,
      "duration": 17.46,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Return partially to the dark apartment with the same orange-shirt protagonist viewing several ambiguous peripheral cues. Foreground: the protagonist notices motion first and turns toward a dark laptop reflection that resembles a person. Midground: a moving curtain, a coat hanging on a chair, and a passerby beyond the glass provide competing explanations. Background: car headlights sweep across the room, creating the actual changing shadow while the protagonist's alert state biases interpretation. Use a branching inference diagram from 'movement detected' to either 'person' or 'passing light,' with the person branch highlighted under high vigilance. The highlighted human silhouette dissolves into an ancestral predator silhouette for scene 14. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Phát hiện trước, giải thích sau\", \"Có người?\", \"Hay ánh đèn?\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"PHÁT HIỆN TRƯỚC, GIẢI THÍCH SAU\", \"CÓ NGƯỜI?\", \"HAY ÁNH ĐÈN?\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "PHÁT HIỆN TRƯỚC, GIẢI THÍCH SAU",
        "CÓ NGƯỜI?",
        "HAY ÁNH ĐÈN?"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_013.png"
    },
    {
      "id": 14,
      "text": "Đây là một phiên bản của cơ chế phát hiện tác nhân hoạt động quá mức, thường được gọi là hyperactive agency detection. Não con người rất giỏi nhận ra tác nhân: sinh vật hoặc con người có ý định, mục tiêu và khả năng hành động. Nhưng hệ thống ấy chấp nhận đánh đổi độ chính xác để lấy tốc độ. Nếu bụi cây rung vì gió mà bạn tưởng là thú săn mồi, cái giá chỉ là một lần giật mình. Nếu bụi cây rung vì thú săn mồi mà bạn tưởng là gió, cái giá có thể lớn hơn nhiều.",
      "start": 266.3,
      "end": 291.68,
      "duration": 25.38,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Move into an evolutionary teaching tableau while preserving the orange-shirt protagonist as an observer walking beside an ancestral stick figure. Foreground: two shaking bushes are compared; one contains only wind and the other conceals a predator. Midground: the agent-detection system rapidly assigns intention and action to an uncertain shape, favoring speed over precision. Background: an evolutionary landscape shows survival consequences, with a harmless false alarm on one branch and severe danger from a missed predator on the other. Use a decision-tree diagram contrasting false positive and false negative costs, with a sensitivity slider set toward caution. Carry the low alarm threshold and alert bell into the modern threat-vigilance scene 15. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Phát hiện tác nhân quá mức\", \"Báo động nhầm\", \"Bỏ sót nguy hiểm\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Phát hiện tác nhân quá mức",
        "Báo động nhầm",
        "Bỏ sót nguy hiểm"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_014.png"
    },
    {
      "id": 15,
      "text": "Threat vigilance, hay sự cảnh giác trước đe dọa, làm ngưỡng báo động hạ thấp hơn nữa. Khi bạn thiếu ngủ, căng thẳng, vừa xem nội dung đáng sợ, đang đi qua nơi xa lạ hoặc từng có trải nghiệm bị theo dõi, não sẽ ưu tiên tín hiệu nguy hiểm. Một tiếng động nhỏ trở nên nổi bật. Một khuôn mặt trung tính có vẻ thiếu thiện cảm hơn. Một cái liếc tình cờ dễ bị diễn giải thành nhìn chằm chằm.",
      "start": 291.68,
      "end": 312.26,
      "duration": 20.58,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Return to a modern urban setting with the orange-shirt protagonist under heightened threat vigilance. Foreground: the tired protagonist holds coffee, carries a stress cloud, and reacts strongly to a tiny noise and a stranger's brief glance. Midground: icons show recent scary media, an unfamiliar street, poor sleep, and a prior following experience lowering the alert threshold. Background: neutral faces appear less friendly and an incidental sideways look is misread as prolonged staring. Use an alarm-threshold gauge dropping from normal to sensitive, with arrows linking each risk factor to amplified interpretation. The gauge becomes a historical loop diagram dated 1898 in scene 16. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Ngưỡng báo động hạ thấp\", \"Thiếu ngủ\", \"Căng thẳng\", \"Caffeine\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"NGƯỠNG BÁO ĐỘNG HẠ THẤP\", \"THIẾU NGỦ\", \"CĂNG THẲNG\", \"CHẤT KÍCH THÍCH\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "NGƯỠNG BÁO ĐỘNG HẠ THẤP",
        "THIẾU NGỦ",
        "CĂNG THẲNG",
        "CHẤT KÍCH THÍCH"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_015.png"
    },
    {
      "id": 16,
      "text": "Năm 1898, nhà tâm lý học Edward Titchener đã viết về “feeling of being stared at”, cảm giác bị nhìn chằm chằm. Ông không xem đây là bằng chứng về một giác quan huyền bí. Titchener đề xuất một vòng phản hồi khá đời thường: bạn bắt đầu nghĩ rằng mình bị nhìn, cơ thể xuất hiện cảm giác khó chịu hoặc căng ở gáy, bạn quay đầu, chuyển động ấy thu hút người phía sau nhìn sang, rồi bạn bắt gặp ánh mắt của họ. Trình tự sau đó bị ghi nhớ như thể ánh mắt của họ đã khiến bạn quay lại ngay từ đầu.",
      "start": 312.26,
      "end": 340.94,
      "duration": 28.68,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Transition into a sepia-accented but still modern educational stick-figure reconstruction of Edward Titchener's 1898 explanation. Foreground: the orange-shirt protagonist thinks about being watched, feels neck tension, turns around, and catches a person behind looking over. Midground: the rear person was initially looking elsewhere but responds to the protagonist's sudden movement. Background: a memory panel later rearranges the sequence as if the rear gaze caused the turn. Use a numbered circular feedback diagram showing thought, bodily discomfort, head turn, attracted gaze, eye contact, and distorted recollection. Break the circle open into a formal experimental trial sequence for scene 17. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"1898\", \"Vòng phản hồi\", \"Ký ức đảo nguyên nhân\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"NĂM 1898\", \"VÒNG PHẢN HỒI\", \"KÝ ỨC ĐỔI THỨ TỰ NGUYÊN NHÂN\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "NĂM 1898",
        "VÒNG PHẢN HỒI",
        "KÝ ỨC ĐỔI THỨ TỰ NGUYÊN NHÂN"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_016.png"
    },
    {
      "id": 17,
      "text": "Một số thí nghiệm của Rupert Sheldrake về khả năng cảm nhận ánh mắt từ phía sau từng được phổ biến rộng rãi và báo cáo kết quả cao hơn ngẫu nhiên trong một số điều kiện. Tuy nhiên, đây vẫn là một lĩnh vực gây tranh cãi. Các nhà phê bình đã nêu vấn đề về thiết kế thí nghiệm, trình tự thử có thể dự đoán, tín hiệu vô tình, cách phân tích và khó khăn khi tái lập kết quả trong điều kiện kiểm soát chặt. Những nghiên cứu ấy có thể được xem như ví dụ về việc một tuyên bố hấp dẫn cần kiểm tra nghiêm ngặt, chứ không phải bằng chứng đã xác lập cho thần giao cách cảm.",
      "start": 340.94,
      "end": 372.58,
      "duration": 31.64,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Set the orange-shirt protagonist inside a controlled experimental review room. Foreground: a participant sits with their back to a possible observer while trial cards alternate between looking and not looking. Midground: one panel presents Rupert Sheldrake's above-chance reports, while another exposes possible predictable sequences, accidental cues, analysis choices, and weak replication under tighter controls. Background: multiple laboratories attempt the same protocol with inconsistent result charts. Use a balance-scale diagram weighing an appealing telepathy claim against design quality, blinding, control, and replication, with the evidence pointer remaining uncertain. The balance resolves into two clearly separated scientific questions in scene 18. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Tuyên bố hấp dẫn\", \"Cần kiểm tra nghiêm ngặt\", \"Khó tái lập\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"TUYÊN BỐ HẤP DẪN\", \"CẦN KIỂM TRA NGHIÊM NGẶT\", \"KHÓ TÁI LẬP\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "TUYÊN BỐ HẤP DẪN",
        "CẦN KIỂM TRA NGHIÊM NGẶT",
        "KHÓ TÁI LẬP"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_017.png"
    },
    {
      "id": 18,
      "text": "Khoa học phải tách hai câu hỏi. Câu thứ nhất là cảm giác bị nhìn có thật hay không. Có, cảm giác chủ quan ấy hoàn toàn thật và đôi khi rất mạnh. Câu thứ hai là cảm giác ấy có chứng minh rằng một người đang nhìn bạn bằng một cơ chế chưa biết hay không. Không. Trải nghiệm chân thật không đồng nghĩa với cách giải thích đầu tiên cũng chân thật.",
      "start": 372.58,
      "end": 391.32,
      "duration": 18.74,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Continue in the review room with the same orange-shirt protagonist between two large question panels. Foreground: the left panel shows a real, intense subjective feeling with body tension and an alert icon; the right panel asks whether that feeling proves an unknown gaze-detection mechanism and shows insufficient evidence. Midground: a check mark validates the experience while a stop line prevents jumping directly to the supernatural explanation. Background: the disputed experiment cards from scene 17 fade into a neutral evidence archive. Use a two-column logic diagram separating experience from explanation, connected only by a dotted rather than causal arrow. The dotted arrow bends into a recalibrating gaze scale for scene 19. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Cảm giác: có thật\", \"Giải thích đầu tiên: chưa chắc đúng\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Cảm giác: có thật",
        "Giải thích đầu tiên: chưa chắc đúng"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_018.png"
    },
    {
      "id": 19,
      "text": "Não còn có thể điều chỉnh chuẩn đánh giá ánh mắt dựa trên kinh nghiệm gần đây. Các nghiên cứu về gaze adaptation, trong đó có công trình của Colin Clifford và các cộng sự, cho thấy sau khi nhìn lâu vào những khuôn mặt có ánh mắt lệch về một hướng, nhận thức về hướng nhìn sau đó có thể bị dịch chuyển. Một ánh mắt vốn hơi lệch có thể trông trực diện hơn, hoặc ngược lại. Nói cách khác, “người đó đang nhìn thẳng vào mình” không phải phép đo tuyệt đối. Đó là kết quả não hiệu chỉnh liên tục theo bối cảnh.",
      "start": 391.32,
      "end": 418.64,
      "duration": 27.32,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Return to the eye-gaze laboratory for an adaptation experiment with the orange-shirt protagonist. Foreground: the protagonist first views a long series of faces looking consistently to one side, then judges a slightly averted test gaze as direct. Midground: a second sequence demonstrates the reverse shift, showing that recent exposure recalibrates the perceived center of gaze. Background: a research card represents Colin Clifford and colleagues, with repeated face arrays documenting gaze adaptation. Use a before-adaptation and after-adaptation dial whose zero point visibly moves, plus arrows from exposure history to changed judgment. The moving zero point becomes an exaggerated self-attention estimate in scene 20. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Thích nghi hướng nhìn\", \"Chuẩn đánh giá bị dịch chuyển\", \"Không phải phép đo tuyệt đối\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Thích nghi hướng nhìn",
        "Chuẩn đánh giá bị dịch chuyển",
        "Không phải phép đo tuyệt đối"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_019.png"
    },
    {
      "id": 20,
      "text": "Rồi còn spotlight effect, hiệu ứng ánh đèn sân khấu. Thomas Gilovich, Victoria Medvec và Kenneth Savitsky đã cho thấy con người thường đánh giá quá cao mức độ người khác chú ý đến ngoại hình và hành vi của mình. Trong một thí nghiệm nổi tiếng, người tham gia mặc một chiếc áo có hình gây ngượng ngùng rồi ước tính có bao nhiêu người trong phòng nhận ra. Họ thường đoán cao hơn thực tế.",
      "start": 418.64,
      "end": 439.94,
      "duration": 21.3,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Stage the spotlight-effect experiment as an educational social scene. Foreground: the orange-shirt protagonist enters a room wearing an intentionally embarrassing graphic shirt and imagines every face turning toward it under a huge spotlight. Midground: in reality, only a few stick figures notice while most continue talking, reading, or looking elsewhere. Background: a result board associated with Thomas Gilovich, Victoria Medvec, and Kenneth Savitsky compares the protagonist's high estimate with the lower observed count. Use a split screen labeled imagined attention versus measured attention, with unequal bar charts and shrinking spotlight cones. The spotlight circle transforms into a phone camera lens for scene 21. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Hiệu ứng ánh đèn sân khấu\", \"Ước tính\", \"Thực tế\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"HIỆU ỨNG ÁNH ĐÈN SÂN KHẤU\", \"ƯỚC TÍNH\", \"THỰC TẾ\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "HIỆU ỨNG ÁNH ĐÈN SÂN KHẤU",
        "ƯỚC TÍNH",
        "THỰC TẾ"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_020.png"
    },
    {
      "id": 21,
      "text": "“Chiếc gương” hiện đại của hiện tượng này là camera trước, cửa kính tối và màn hình điện thoại. Bạn bước vào thang máy, thấy một người trên màn hình an ninh đang nhìn mình, rồi nhận ra đó là hình ảnh của chính bạn với độ trễ nhỏ. Bạn lướt mạng xã hội và biết lượt xem, trạng thái trực tuyến, dấu “đã đọc” đều có thể được ghi lại. Môi trường số khiến ý niệm bị quan sát luôn hiện diện, ngay cả khi không có đôi mắt nào trong phòng. Não cổ xưa đang xử lý một thế giới đầy camera, chỉ báo hoạt động và hình ảnh phản chiếu.",
      "start": 439.94,
      "end": 468.3,
      "duration": 28.36,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Move into a dense digital-surveillance montage while keeping the orange-shirt protagonist central. Foreground: in an elevator, the protagonist sees a security-screen figure staring back, then recognizes it as their own slightly delayed image. Midground: a dark glass reflection, front-facing phone camera, view counter, online-status dot, and read receipt surround the protagonist with persistent observation cues. Background: a city network of cameras and screens contrasts an ancient brain icon with the modern monitored environment. Use a multi-panel mirror diagram connecting physical reflection, recorded image, activity indicator, and imagined observer, with delay arrows clarifying the elevator illusion. Collapse the panels into a practical evidence-check interface for scene 22. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Lượt xem\", \"Đang hoạt động\", \"Đã đọc\", \"Não cổ xưa, thế giới đầy camera\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"LƯỢT XEM\", \"ĐANG HOẠT ĐỘNG\", \"ĐÃ ĐỌC\", \"NÃO CỔ XƯA, THẾ GIỚI ĐẦY CAMERA\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device. Use exactly four readable labels total and no more: “LƯỢT XEM”, “ĐANG HOẠT ĐỘNG”, “ĐÃ ĐỌC”, “NÃO CỔ XƯA, THẾ GIỚI ĐẦY CAMERA”. All other interface elements must be icon-only solid shapes without letters, captions, labels, numbers, or pseudo-text.",
      "visible_text": [
        "LƯỢT XEM",
        "ĐANG HOẠT ĐỘNG",
        "ĐÃ ĐỌC",
        "NÃO CỔ XƯA, THẾ GIỚI ĐẦY CAMERA"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_021.png"
    },
    {
      "id": 22,
      "text": "Khi cảm giác bị theo dõi xuất hiện, bước đầu tiên không phải phủ nhận nó, cũng không phải tin tuyệt đối vào nó. Hãy gọi đúng tên trải nghiệm: “Mình đang có cảm giác bị quan sát.” Câu này tách cảm giác khỏi kết luận “Có người chắc chắn đang theo dõi mình.” Sau đó, kiểm tra bằng chứng có thể quan sát được: có người thật sự đổi hướng theo bạn không, có cùng một người xuất hiện lặp lại qua nhiều địa điểm không, có tiếng bước chân đồng bộ, tin nhắn đe dọa, camera lạ hoặc dấu hiệu xâm nhập không?",
      "start": 468.3,
      "end": 494.88,
      "duration": 26.58,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Shift to a calm safety-assessment interface with the orange-shirt protagonist grounding themself. Foreground: the protagonist places the feeling into a labeled thought bubble rather than declaring that a watcher certainly exists. Midground: an evidence checklist examines whether someone repeatedly changes direction with them, appears across multiple locations, produces synchronized footsteps, sends threats, or has placed an unfamiliar camera. Background: the digital-surveillance montage from scene 21 fades into individually testable observations rather than one ominous cloud. Use a two-lane flowchart separating 'I feel observed' from 'observable evidence,' with a magnifying-glass arrow inspecting each concrete sign. Continue from the checklist into false-alarm reduction steps in scene 23. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Mình đang có cảm giác bị quan sát\", \"Kiểm tra bằng chứng\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Mình đang có cảm giác bị quan sát",
        "Kiểm tra bằng chứng"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_022.png"
    },
    {
      "id": 23,
      "text": "Bạn cũng có thể kiểm tra các yếu tố làm tăng sai báo động: thiếu ngủ, caffeine, căng thẳng, bóng phản chiếu, âm thanh từ thiết bị, nội dung vừa xem và tầm nhìn bị che khuất. Thở chậm hơn, bật thêm đèn, tháo tai nghe và nhìn toàn cảnh thay vì khóa chú ý vào một chi tiết. Nếu đang ở nơi công cộng, hãy di chuyển sang vị trí sáng, đông người và dễ quan sát. Nếu cảm giác lặp lại thường xuyên, gây mất ngủ hoặc xuất hiện dù không có bằng chứng trong nhiều bối cảnh, việc trao đổi với chuyên gia sức khỏe tâm thần có thể giúp giảm gánh nặng cảnh giác kéo dài.",
      "start": 494.88,
      "end": 526.36,
      "duration": 31.48,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Continue the practical sequence with the same orange-shirt protagonist moving through a calming checklist. Foreground: the protagonist breathes slowly, switches on more lights, removes headphones, and widens their gaze instead of fixating on one shadow. Midground: icons identify sleep loss, caffeine, stress, reflections, device sounds, frightening content, and obstructed vision as false-alarm amplifiers. Background: an alternate public-space panel shows the protagonist moving toward a bright, populated, observable area; another small panel shows a supportive mental-health professional for persistent, sleep-disrupting episodes without evidence. Use a before-and-after arousal gauge and a panoramic scan arrow. A concrete danger marker remains highlighted and leads into the real-safety branch of scene 24. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Thở chậm\", \"Bật đèn\", \"Tháo tai nghe\", \"Nhìn toàn cảnh\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"1. THỞ CHẬM\", \"2. BẬT ĐÈN\", \"3. THÁO TAI NGHE\", \"4. NHÌN TOÀN CẢNH\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "1. THỞ CHẬM",
        "2. BẬT ĐÈN",
        "3. THÁO TAI NGHE",
        "4. NHÌN TOÀN CẢNH"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_023.png"
    },
    {
      "id": 24,
      "text": "Nhưng phân biệt cảm giác với bằng chứng không có nghĩa là bỏ qua an toàn. Nếu có dấu hiệu cụ thể cho thấy bạn đang bị bám theo hoặc gặp nguy hiểm, hãy ưu tiên rời khỏi nơi vắng, vào cửa hàng hay khu vực có nhân viên an ninh, gọi người đáng tin cậy, chia sẻ vị trí và liên hệ dịch vụ khẩn cấp tại địa phương. Đừng quay về nhà nếu điều đó có thể tiết lộ nơi ở. Trong tình huống thực tế, an toàn quan trọng hơn việc chứng minh cảm giác đúng hay sai.",
      "start": 526.36,
      "end": 551.04,
      "duration": 24.68,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Present a decisive real-world safety flowchart without losing the educational stick-figure style. Foreground: after noticing concrete signs of being followed, the orange-shirt protagonist moves away from an isolated route and enters a bright store with staff and security. Midground: the protagonist calls a trusted person, shares live location, and prepares to contact local emergency services. Background: a map shows the route deliberately avoiding home so the residence is not revealed, while a follower silhouette stays at a safe distance. Use a red-to-green decision diagram prioritizing leave, seek people, contact support, share location, and avoid going home. End with the protagonist safely returning to the original apartment only after the instructional sequence, setting up scene 25. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Ưu tiên an toàn\", \"Đến nơi sáng, đông người\", \"Không quay về nhà\", \"Gọi hỗ trợ\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style.",
      "visible_text": [
        "Ưu tiên an toàn",
        "Đến nơi sáng, đông người",
        "Không quay về nhà",
        "Gọi hỗ trợ"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_024.png"
    },
    {
      "id": 25,
      "text": "Và rồi bạn trở lại căn phòng ban đầu: laptop vẫn sáng, tai nghe vẫn nằm trên bàn, cửa sổ vẫn phản chiếu một hình người. Lần này, khi cảm giác có ai đang nhìn xuất hiện, bạn không cần chế giễu nó và cũng không cần biến nó thành năng lực siêu nhiên. Có thể ngoài kia thật sự có một ánh mắt, nên bạn bình tĩnh kiểm tra. Cũng có thể tầm nhìn ngoại vi vừa bắt được bóng của chính bạn, amygdala đã đánh dấu nó là quan trọng, STS tìm kiếm tín hiệu xã hội, còn hệ cảnh giác dựng nên một tác nhân để giải thích khoảng trống.",
      "start": 551.04,
      "end": 580.28,
      "duration": 29.24,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Return precisely to the opening dark apartment with the recurring white round-headed protagonist in the orange shirt. Foreground: the protagonist sits calmly beside the glowing laptop, with headphones now on the table, and notices a human-shaped reflection in the window before checking it methodically. Midground: layered transparent icons show peripheral vision detecting the shadow, the amygdala flagging importance, STS searching for social gaze, and the vigilance system proposing an agent. Background: the locked door, still curtain, and sparse illuminated windows mirror scenes 1 and 2, while a possible real observer and the protagonist's own reflection remain separate hypotheses. Use a branching brain-to-evidence diagram ending in 'check calmly' rather than panic or dismissal. Fade the scientific layers into the reflective final image of scene 26. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Bình tĩnh kiểm tra\", \"Ngoại vi\", \"Amygdala\", \"STS\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"BÌNH TĨNH KIỂM TRA\", \"TẦM NHÌN NGOẠI VI\", \"HẠCH HẠNH NHÂN\", \"RÃNH THÁI DƯƠNG TRÊN\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "BÌNH TĨNH KIỂM TRA",
        "TẦM NHÌN NGOẠI VI",
        "HẠCH HẠNH NHÂN",
        "RÃNH THÁI DƯƠNG TRÊN"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_025.png"
    },
    {
      "id": 26,
      "text": "Căn phòng có thể vẫn trống, nhưng cảm giác không đến từ hư vô. Nó đến từ một bộ não được xây dựng để phát hiện ánh mắt, dự đoán ý định và thà báo động nhầm còn hơn im lặng trước nguy hiểm. Điều kỳ lạ cuối cùng không phải là bạn có thể cảm thấy một người đang nhìn khi chẳng có ai. Điều kỳ lạ là trong bóng tối, đôi khi chính bộ não đang nhìn bạn nhìn lại thế giới, rồi khiến bạn tưởng thế giới đang nhìn lại mình.",
      "start": 580.28,
      "end": 603.26,
      "duration": 22.98,
      "prompt": "Hand-drawn 2D doodle cartoon animation, flat colors, bold black outlines, slightly imperfect sketchy marker lines, simple expressive stick-figure characters with large circular white heads, dot eyes, thick eyebrow lines, information-rich educational doodle scene, layered visual storytelling, flat solid color blocks, bold black marker outlines, no storybook painting, no realistic anatomy, no paper texture. Recurring protagonist: large circular white head, dot eyes, thick expressive eyebrows, narrow orange shirt torso, stick limbs, consistent proportions. Close in the same apartment with a quiet, information-rich visual synthesis. Foreground: the orange-shirt protagonist stands before the window, facing both the dark room and their own reflection, no longer startled but attentively evaluating the scene. Midground: a transparent brain contains three linked systems for detecting gaze, predicting intention, and favoring false alarms over missed danger; faint arrows connect them to the neck tension and watcher thought from scene 1. Background: the apartment remains empty while the city lights resemble scattered eyes only until the diagram resolves them into ordinary windows. Use a mirrored split-screen in which the brain observes the world on one side and the world seems to observe the protagonist on the other, then merge both halves into one calm frame. Echo the opening composition exactly, completing the apartment-to-lab-to-evolution-to-safety-to-apartment continuity. Use channel palette #F5820D #2D5FBF #3A9E3A #F5C518 #D94040 #8B5E3C #6EB5E8 #C4965A #FFFFFF. Intentional visible Vietnamese text, render exactly with correct diacritics and no other words: [\"Phát hiện ánh mắt\", \"Dự đoán ý định\", \"Thà báo động nhầm\"]. Use only clean uniform fill regions from the fixed palette; every wall, window, object, screen, shirt, face, and silhouette must be one solid color with hard boundaries. No glow, light rays, haze, reflection blur, transparency, grain, noise, speckles, surface variation, highlights, bevels, soft edges, color ramps, gradients, shadows, textures, photorealism, or 3D, 16:9 aspect ratio, educational YouTube explainer doodle style. FINAL TEXT OVERRIDE: The only readable words anywhere are exactly [\"PHÁT HIỆN ÁNH MẮT\", \"DỰ ĐOÁN Ý ĐỊNH\", \"THÀ BÁO ĐỘNG NHẦM\"]. No other letters, labels, English words, pseudo-text, timestamps, or watermarks. Keep all people as simple stick figures with circular white heads, dot eyes, thick brows, narrow flat torsos, and stick limbs; no hair, ears, realistic faces, rounded full bodies, cinematic lighting, wood grain, or surface texture. Preserve a clear focal hierarchy while retaining the narration-relevant setting, props, and explanatory device.",
      "visible_text": [
        "PHÁT HIỆN ÁNH MẮT",
        "DỰ ĐOÁN Ý ĐỊNH",
        "THÀ BÁO ĐỘNG NHẦM"
      ],
      "output_path": "/data/video-pipeline/StickerMan/project/003-Vi-Sao-Nao-Bo-Khien-Ban-Cam-Thay-Co-Nguoi-Dang-Nhin-Minh/images/scenes/scene_026.png"
    }
  ]
}
