[{"data":1,"prerenderedAt":1861},["ShallowReactive",2],{"page-\u002Fprompt-engineering\u002F11-self-consistency-and-verification":3},{"id":4,"title":5,"body":6,"description":1854,"extension":1855,"meta":1856,"navigation":70,"path":1857,"seo":1858,"stem":1859,"__hash__":1860},"content\u002Fprompt-engineering\u002F11-self-consistency-and-verification.md","11 — Self-Consistency & Verification",{"type":7,"value":8,"toc":1842},"minimark",[9,13,18,155,159,223,264,268,751,755,989,993,1324,1328,1450,1454,1563,1567,1676,1680,1684,1728,1732,1838],[10,11,5],"h1",{"id":12},"_11-self-consistency-verification",[14,15,17],"h2",{"id":16},"why-naive-self-verification-fails","Why Naive Self-Verification Fails",[19,20,23],"code-wrapper",{"filename":21,"language":22},"naive_self_verification.py","python",[24,25,29],"pre",{"className":26,"code":27,"language":22,"meta":28,"style":28},"language-python shiki shiki-themes github-light github-dark","# NAIVE: \"double-check your answer\" or \"are you sure?\"\n# The same model, conditioned on the SAME reasoning, re-examines its OWN output.\n# If the error came from a genuine gap in what the model learned (not a sampling\n# fluke), re-examining the same reasoning reproduces the SAME error with renewed\n# confidence.\n\n# EXAMPLE: \"What is the capital of Australia?\"\n# Model answers \"Sydney\" (a common misconception — Sydney is the best-known\n# city but NOT the capital; it's Canberra).\n# \"Are you sure?\" → either \"Yes, I'm sure\" (no re-derivation) or flips to a\n# DIFFERENT wrong answer because the question implied doubt, not because new\n# evidence was introduced.\n\n# NAIVE SELF-VERIFICATION WORKS FOR:\n# - Arithmetic slips visible from the output itself (re-read catches them)\n# - JSON with a missing bracket\n# - Answers contradicting an explicit constraint stated earlier in the same prompt\n# It FAILS FOR:\n# - Factual errors rooted in what the model BELIEVES to be true\n# - No new signal is introduced that would change the belief on a second pass\n","",[30,31,32,41,47,53,59,65,72,78,84,90,96,102,108,113,119,125,131,137,143,149],"code",{"__ignoreMap":28},[33,34,37],"span",{"class":35,"line":36},"line",1,[33,38,40],{"class":39},"sdCPZ","# NAIVE: \"double-check your answer\" or \"are you sure?\"\n",[33,42,44],{"class":35,"line":43},2,[33,45,46],{"class":39},"# The same model, conditioned on the SAME reasoning, re-examines its OWN output.\n",[33,48,50],{"class":35,"line":49},3,[33,51,52],{"class":39},"# If the error came from a genuine gap in what the model learned (not a sampling\n",[33,54,56],{"class":35,"line":55},4,[33,57,58],{"class":39},"# fluke), re-examining the same reasoning reproduces the SAME error with renewed\n",[33,60,62],{"class":35,"line":61},5,[33,63,64],{"class":39},"# confidence.\n",[33,66,68],{"class":35,"line":67},6,[33,69,71],{"emptyLinePlaceholder":70},true,"\n",[33,73,75],{"class":35,"line":74},7,[33,76,77],{"class":39},"# EXAMPLE: \"What is the capital of Australia?\"\n",[33,79,81],{"class":35,"line":80},8,[33,82,83],{"class":39},"# Model answers \"Sydney\" (a common misconception — Sydney is the best-known\n",[33,85,87],{"class":35,"line":86},9,[33,88,89],{"class":39},"# city but NOT the capital; it's Canberra).\n",[33,91,93],{"class":35,"line":92},10,[33,94,95],{"class":39},"# \"Are you sure?\" → either \"Yes, I'm sure\" (no re-derivation) or flips to a\n",[33,97,99],{"class":35,"line":98},11,[33,100,101],{"class":39},"# DIFFERENT wrong answer because the question implied doubt, not because new\n",[33,103,105],{"class":35,"line":104},12,[33,106,107],{"class":39},"# evidence was introduced.\n",[33,109,111],{"class":35,"line":110},13,[33,112,71],{"emptyLinePlaceholder":70},[33,114,116],{"class":35,"line":115},14,[33,117,118],{"class":39},"# NAIVE SELF-VERIFICATION WORKS FOR:\n",[33,120,122],{"class":35,"line":121},15,[33,123,124],{"class":39},"# - Arithmetic slips visible from the output itself (re-read catches them)\n",[33,126,128],{"class":35,"line":127},16,[33,129,130],{"class":39},"# - JSON with a missing bracket\n",[33,132,134],{"class":35,"line":133},17,[33,135,136],{"class":39},"# - Answers contradicting an explicit constraint stated earlier in the same prompt\n",[33,138,140],{"class":35,"line":139},18,[33,141,142],{"class":39},"# It FAILS FOR:\n",[33,144,146],{"class":35,"line":145},19,[33,147,148],{"class":39},"# - Factual errors rooted in what the model BELIEVES to be true\n",[33,150,152],{"class":35,"line":151},20,[33,153,154],{"class":39},"# - No new signal is introduced that would change the belief on a second pass\n",[14,156,158],{"id":157},"structured-verification-different-method-not-re-read","Structured Verification: Different Method, Not Re-Read",[19,160,163],{"filename":161,"language":162},"structured_verification.md","markdown",[24,164,167],{"className":165,"code":166,"language":162,"meta":28,"style":28},"language-markdown shiki shiki-themes github-light github-dark","Step 1: Solve the following problem, showing your work.\n\nA store offers a 20% discount, then charges 8% sales tax on the\ndiscounted price. If the original price is $150, what is the final price?\n\nStep 2: Now verify your Step 1 answer by solving the problem a second\ntime using a different method (e.g., if you multiplied discount and tax\nfactors together first, this time apply the discount first, get an\nintermediate dollar value, then apply tax to that intermediate value\nseparately). State whether the two methods agree. If they disagree,\nidentify which step diverged and redo the calculation.\n",[30,168,169,175,179,184,189,193,198,203,208,213,218],{"__ignoreMap":28},[33,170,171],{"class":35,"line":36},[33,172,174],{"class":173},"ssxIu","Step 1: Solve the following problem, showing your work.\n",[33,176,177],{"class":35,"line":43},[33,178,71],{"emptyLinePlaceholder":70},[33,180,181],{"class":35,"line":49},[33,182,183],{"class":173},"A store offers a 20% discount, then charges 8% sales tax on the\n",[33,185,186],{"class":35,"line":55},[33,187,188],{"class":173},"discounted price. If the original price is $150, what is the final price?\n",[33,190,191],{"class":35,"line":61},[33,192,71],{"emptyLinePlaceholder":70},[33,194,195],{"class":35,"line":67},[33,196,197],{"class":173},"Step 2: Now verify your Step 1 answer by solving the problem a second\n",[33,199,200],{"class":35,"line":74},[33,201,202],{"class":173},"time using a different method (e.g., if you multiplied discount and tax\n",[33,204,205],{"class":35,"line":80},[33,206,207],{"class":173},"factors together first, this time apply the discount first, get an\n",[33,209,210],{"class":35,"line":86},[33,211,212],{"class":173},"intermediate dollar value, then apply tax to that intermediate value\n",[33,214,215],{"class":35,"line":92},[33,216,217],{"class":173},"separately). State whether the two methods agree. If they disagree,\n",[33,219,220],{"class":35,"line":98},[33,221,222],{"class":173},"identify which step diverged and redo the calculation.\n",[19,224,226],{"filename":225,"language":22},"verification_principle.py",[24,227,229],{"className":26,"code":228,"language":22,"meta":28,"style":28},"# PRINCIPLE: verification is only informative when it's STRUCTURALLY DIFFERENT\n# from the original reasoning path, not just a repetition phrased as a question.\n\n# BAD: \"double-check your work\" → likely re-reads the same derivation and nods\n# GOOD: \"solve using a different method and compare\" → genuinely independent derivation\n\n# The different method is real corroborating evidence. A re-read is not.\n",[30,230,231,236,241,245,250,255,259],{"__ignoreMap":28},[33,232,233],{"class":35,"line":36},[33,234,235],{"class":39},"# PRINCIPLE: verification is only informative when it's STRUCTURALLY DIFFERENT\n",[33,237,238],{"class":35,"line":43},[33,239,240],{"class":39},"# from the original reasoning path, not just a repetition phrased as a question.\n",[33,242,243],{"class":35,"line":49},[33,244,71],{"emptyLinePlaceholder":70},[33,246,247],{"class":35,"line":55},[33,248,249],{"class":39},"# BAD: \"double-check your work\" → likely re-reads the same derivation and nods\n",[33,251,252],{"class":35,"line":61},[33,253,254],{"class":39},"# GOOD: \"solve using a different method and compare\" → genuinely independent derivation\n",[33,256,257],{"class":35,"line":67},[33,258,71],{"emptyLinePlaceholder":70},[33,260,261],{"class":35,"line":74},[33,262,263],{"class":39},"# The different method is real corroborating evidence. A re-read is not.\n",[14,265,267],{"id":266},"self-consistency-sample-n-times-vote","Self-Consistency: Sample N Times, Vote",[19,269,271],{"filename":270,"language":22},"self_consistency.py",[24,272,274],{"className":26,"code":273,"language":22,"meta":28,"style":28},"from collections import Counter\nimport re\n\ndef self_consistency(prompt: str, n: int = 5, temperature: float = 0.7) -> tuple[str, float]:\n    \"\"\"Generate N independent reasoning paths at nonzero temperature, take majority vote.\n\n    Intuition: a model's reasoning errors on a hard problem are often RANDOM with\n    respect to which wrong path it takes. Correct reasoning tends to converge on\n    the same answer via multiple valid routes. Many ways to be wrong, usually one\n    way to be right → the correct answer shows up more often than any single wrong one.\n    \"\"\"\n    answers = []\n    for _ in range(n):\n        response = client.messages.create(\n            model=\"claude-opus-5\",\n            max_tokens=1024,\n            temperature=temperature,  # CRITICAL: must be nonzero for sampling diversity\n            messages=[{\"role\": \"user\", \"content\": prompt}],\n        )\n        answer = extract_final_answer(response.content[0].text)\n        answers.append(answer)\n\n    # Majority vote\n    most_common, count = Counter(answers).most_common(1)[0]\n    confidence = count \u002F n\n    return most_common, confidence\n\ndef extract_final_answer(text: str) -> str:\n    \"\"\"Extract the final answer from a CoT response.\"\"\"\n    # Look for \"The answer is X\" or \"Final answer: X\" patterns\n    match = re.search(r'(?:answer is|Final answer:)\\s*(.+?)(?:\\n|$)', text, re.IGNORECASE)\n    return match.group(1).strip() if match else text.strip().split('\\n')[-1]\n\n# WHAT IT FIXES: variance — cases where the model sometimes gets it right\n# WHAT IT DOESN'T FIX: bias — systematic errors the model makes EVERY time\n# (the Sydney\u002FCanberra case would show \"Sydney\" across ALL n samples —\n#  it's not a random slip, it's a consistent misconception)\n",[30,275,276,291,298,302,354,360,364,369,374,379,384,389,400,417,427,441,453,466,493,498,514,520,525,531,553,570,579,584,605,611,617,682,722,727,733,739,745],{"__ignoreMap":28},[33,277,278,282,285,288],{"class":35,"line":36},[33,279,281],{"class":280},"svdQ7","from",[33,283,284],{"class":173}," collections ",[33,286,287],{"class":280},"import",[33,289,290],{"class":173}," Counter\n",[33,292,293,295],{"class":35,"line":43},[33,294,287],{"class":280},[33,296,297],{"class":173}," re\n",[33,299,300],{"class":35,"line":49},[33,301,71],{"emptyLinePlaceholder":70},[33,303,304,307,311,314,318,321,324,327,330,333,336,338,341,344,346,349,351],{"class":35,"line":55},[33,305,306],{"class":280},"def",[33,308,310],{"class":309},"sIsaT"," self_consistency",[33,312,313],{"class":173},"(prompt: ",[33,315,317],{"class":316},"snvgF","str",[33,319,320],{"class":173},", n: ",[33,322,323],{"class":316},"int",[33,325,326],{"class":280}," =",[33,328,329],{"class":316}," 5",[33,331,332],{"class":173},", temperature: ",[33,334,335],{"class":316},"float",[33,337,326],{"class":280},[33,339,340],{"class":316}," 0.7",[33,342,343],{"class":173},") -> tuple[",[33,345,317],{"class":316},[33,347,348],{"class":173},", ",[33,350,335],{"class":316},[33,352,353],{"class":173},"]:\n",[33,355,356],{"class":35,"line":61},[33,357,359],{"class":358},"sJ6F3","    \"\"\"Generate N independent reasoning paths at nonzero temperature, take majority vote.\n",[33,361,362],{"class":35,"line":67},[33,363,71],{"emptyLinePlaceholder":70},[33,365,366],{"class":35,"line":74},[33,367,368],{"class":358},"    Intuition: a model's reasoning errors on a hard problem are often RANDOM with\n",[33,370,371],{"class":35,"line":80},[33,372,373],{"class":358},"    respect to which wrong path it takes. Correct reasoning tends to converge on\n",[33,375,376],{"class":35,"line":86},[33,377,378],{"class":358},"    the same answer via multiple valid routes. Many ways to be wrong, usually one\n",[33,380,381],{"class":35,"line":92},[33,382,383],{"class":358},"    way to be right → the correct answer shows up more often than any single wrong one.\n",[33,385,386],{"class":35,"line":98},[33,387,388],{"class":358},"    \"\"\"\n",[33,390,391,394,397],{"class":35,"line":104},[33,392,393],{"class":173},"    answers ",[33,395,396],{"class":280},"=",[33,398,399],{"class":173}," []\n",[33,401,402,405,408,411,414],{"class":35,"line":110},[33,403,404],{"class":280},"    for",[33,406,407],{"class":173}," _ ",[33,409,410],{"class":280},"in",[33,412,413],{"class":316}," range",[33,415,416],{"class":173},"(n):\n",[33,418,419,422,424],{"class":35,"line":115},[33,420,421],{"class":173},"        response ",[33,423,396],{"class":280},[33,425,426],{"class":173}," client.messages.create(\n",[33,428,429,433,435,438],{"class":35,"line":121},[33,430,432],{"class":431},"sCrzJ","            model",[33,434,396],{"class":280},[33,436,437],{"class":358},"\"claude-opus-5\"",[33,439,440],{"class":173},",\n",[33,442,443,446,448,451],{"class":35,"line":127},[33,444,445],{"class":431},"            max_tokens",[33,447,396],{"class":280},[33,449,450],{"class":316},"1024",[33,452,440],{"class":173},[33,454,455,458,460,463],{"class":35,"line":133},[33,456,457],{"class":431},"            temperature",[33,459,396],{"class":280},[33,461,462],{"class":173},"temperature,  ",[33,464,465],{"class":39},"# CRITICAL: must be nonzero for sampling diversity\n",[33,467,468,471,473,476,479,482,485,487,490],{"class":35,"line":139},[33,469,470],{"class":431},"            messages",[33,472,396],{"class":280},[33,474,475],{"class":173},"[{",[33,477,478],{"class":358},"\"role\"",[33,480,481],{"class":173},": ",[33,483,484],{"class":358},"\"user\"",[33,486,348],{"class":173},[33,488,489],{"class":358},"\"content\"",[33,491,492],{"class":173},": prompt}],\n",[33,494,495],{"class":35,"line":145},[33,496,497],{"class":173},"        )\n",[33,499,500,503,505,508,511],{"class":35,"line":151},[33,501,502],{"class":173},"        answer ",[33,504,396],{"class":280},[33,506,507],{"class":173}," extract_final_answer(response.content[",[33,509,510],{"class":316},"0",[33,512,513],{"class":173},"].text)\n",[33,515,517],{"class":35,"line":516},21,[33,518,519],{"class":173},"        answers.append(answer)\n",[33,521,523],{"class":35,"line":522},22,[33,524,71],{"emptyLinePlaceholder":70},[33,526,528],{"class":35,"line":527},23,[33,529,530],{"class":39},"    # Majority vote\n",[33,532,534,537,539,542,545,548,550],{"class":35,"line":533},24,[33,535,536],{"class":173},"    most_common, count ",[33,538,396],{"class":280},[33,540,541],{"class":173}," Counter(answers).most_common(",[33,543,544],{"class":316},"1",[33,546,547],{"class":173},")[",[33,549,510],{"class":316},[33,551,552],{"class":173},"]\n",[33,554,556,559,561,564,567],{"class":35,"line":555},25,[33,557,558],{"class":173},"    confidence ",[33,560,396],{"class":280},[33,562,563],{"class":173}," count ",[33,565,566],{"class":280},"\u002F",[33,568,569],{"class":173}," n\n",[33,571,573,576],{"class":35,"line":572},26,[33,574,575],{"class":280},"    return",[33,577,578],{"class":173}," most_common, confidence\n",[33,580,582],{"class":35,"line":581},27,[33,583,71],{"emptyLinePlaceholder":70},[33,585,587,589,592,595,597,600,602],{"class":35,"line":586},28,[33,588,306],{"class":280},[33,590,591],{"class":309}," extract_final_answer",[33,593,594],{"class":173},"(text: ",[33,596,317],{"class":316},[33,598,599],{"class":173},") -> ",[33,601,317],{"class":316},[33,603,604],{"class":173},":\n",[33,606,608],{"class":35,"line":607},29,[33,609,610],{"class":358},"    \"\"\"Extract the final answer from a CoT response.\"\"\"\n",[33,612,614],{"class":35,"line":613},30,[33,615,616],{"class":39},"    # Look for \"The answer is X\" or \"Final answer: X\" patterns\n",[33,618,620,623,625,628,631,634,637,641,644,647,650,653,656,659,662,666,668,671,673,676,679],{"class":35,"line":619},31,[33,621,622],{"class":173},"    match ",[33,624,396],{"class":280},[33,626,627],{"class":173}," re.search(",[33,629,630],{"class":280},"r",[33,632,633],{"class":358},"'",[33,635,636],{"class":316},"(?:",[33,638,640],{"class":639},"svAP2","answer is",[33,642,643],{"class":280},"|",[33,645,646],{"class":639},"Final answer:",[33,648,649],{"class":316},")\\s",[33,651,652],{"class":280},"*",[33,654,655],{"class":316},"(.",[33,657,658],{"class":280},"+?",[33,660,661],{"class":316},")(?:",[33,663,665],{"class":664},"snRuI","\\n",[33,667,643],{"class":280},[33,669,670],{"class":316},"$)",[33,672,633],{"class":358},[33,674,675],{"class":173},", text, re.",[33,677,678],{"class":316},"IGNORECASE",[33,680,681],{"class":173},")\n",[33,683,685,687,690,692,695,698,701,704,707,709,711,713,715,718,720],{"class":35,"line":684},32,[33,686,575],{"class":280},[33,688,689],{"class":173}," match.group(",[33,691,544],{"class":316},[33,693,694],{"class":173},").strip() ",[33,696,697],{"class":280},"if",[33,699,700],{"class":173}," match ",[33,702,703],{"class":280},"else",[33,705,706],{"class":173}," text.strip().split(",[33,708,633],{"class":358},[33,710,665],{"class":316},[33,712,633],{"class":358},[33,714,547],{"class":173},[33,716,717],{"class":280},"-",[33,719,544],{"class":316},[33,721,552],{"class":173},[33,723,725],{"class":35,"line":724},33,[33,726,71],{"emptyLinePlaceholder":70},[33,728,730],{"class":35,"line":729},34,[33,731,732],{"class":39},"# WHAT IT FIXES: variance — cases where the model sometimes gets it right\n",[33,734,736],{"class":35,"line":735},35,[33,737,738],{"class":39},"# WHAT IT DOESN'T FIX: bias — systematic errors the model makes EVERY time\n",[33,740,742],{"class":35,"line":741},36,[33,743,744],{"class":39},"# (the Sydney\u002FCanberra case would show \"Sydney\" across ALL n samples —\n",[33,746,748],{"class":35,"line":747},37,[33,749,750],{"class":39},"#  it's not a random slip, it's a consistent misconception)\n",[14,752,754],{"id":753},"verification-against-external-ground-truth","Verification Against External Ground Truth",[19,756,758],{"filename":757,"language":22},"external_verification.py",[24,759,761],{"className":26,"code":760,"language":22,"meta":28,"style":28},"def verified_extraction(document: str, schema: dict) -> dict:\n    \"\"\"Extract structured data, then verify extracted quotes against source text.\n    This is a DETERMINISTIC, non-LLM check — cheaper and more reliable than asking\n    the model to check itself.\"\"\"\n    result = extract_with_schema(document, schema)\n\n    for field, value in result.items():\n        if value is not None and isinstance(value, str):\n            # Check: does this extracted value actually appear in the source?\n            if value not in document:\n                flag_for_review(\n                    field=field,\n                    value=value,\n                    reason=\"Extracted value not found verbatim in source — likely fabricated or paraphrased\"\n                )\n\n    return result\n\n# This catches a REAL, common failure mode: a model paraphrasing or subtly\n# fabricating a \"quote\" that doesn't appear in the source. Far more reliable than\n# asking the model \"is this quote accurate?\" (same model checking its own output).\n\n# GENERAL PRINCIPLE: verification is STRONGEST when checking against ground truth\n# EXTERNAL to the model's generation:\n#   - a database (does this customer ID exist?)\n#   - a calculator (is this arithmetic result correct?)\n#   - the literal source document (does this quote appear verbatim?)\n#   - a schema validator (is this JSON structurally valid?)\n# All of these are cheap, deterministic, and don't depend on the model grading itself.\n",[30,762,763,787,792,797,802,812,816,828,859,864,880,885,895,905,915,920,924,931,935,940,945,950,954,959,964,969,974,979,984],{"__ignoreMap":28},[33,764,765,767,770,773,775,778,781,783,785],{"class":35,"line":36},[33,766,306],{"class":280},[33,768,769],{"class":309}," verified_extraction",[33,771,772],{"class":173},"(document: ",[33,774,317],{"class":316},[33,776,777],{"class":173},", schema: ",[33,779,780],{"class":316},"dict",[33,782,599],{"class":173},[33,784,780],{"class":316},[33,786,604],{"class":173},[33,788,789],{"class":35,"line":43},[33,790,791],{"class":358},"    \"\"\"Extract structured data, then verify extracted quotes against source text.\n",[33,793,794],{"class":35,"line":49},[33,795,796],{"class":358},"    This is a DETERMINISTIC, non-LLM check — cheaper and more reliable than asking\n",[33,798,799],{"class":35,"line":55},[33,800,801],{"class":358},"    the model to check itself.\"\"\"\n",[33,803,804,807,809],{"class":35,"line":61},[33,805,806],{"class":173},"    result ",[33,808,396],{"class":280},[33,810,811],{"class":173}," extract_with_schema(document, schema)\n",[33,813,814],{"class":35,"line":67},[33,815,71],{"emptyLinePlaceholder":70},[33,817,818,820,823,825],{"class":35,"line":74},[33,819,404],{"class":280},[33,821,822],{"class":173}," field, value ",[33,824,410],{"class":280},[33,826,827],{"class":173}," result.items():\n",[33,829,830,833,836,839,842,845,848,851,854,856],{"class":35,"line":80},[33,831,832],{"class":280},"        if",[33,834,835],{"class":173}," value ",[33,837,838],{"class":280},"is",[33,840,841],{"class":280}," not",[33,843,844],{"class":316}," None",[33,846,847],{"class":280}," and",[33,849,850],{"class":316}," isinstance",[33,852,853],{"class":173},"(value, ",[33,855,317],{"class":316},[33,857,858],{"class":173},"):\n",[33,860,861],{"class":35,"line":86},[33,862,863],{"class":39},"            # Check: does this extracted value actually appear in the source?\n",[33,865,866,869,871,874,877],{"class":35,"line":92},[33,867,868],{"class":280},"            if",[33,870,835],{"class":173},[33,872,873],{"class":280},"not",[33,875,876],{"class":280}," in",[33,878,879],{"class":173}," document:\n",[33,881,882],{"class":35,"line":98},[33,883,884],{"class":173},"                flag_for_review(\n",[33,886,887,890,892],{"class":35,"line":104},[33,888,889],{"class":431},"                    field",[33,891,396],{"class":280},[33,893,894],{"class":173},"field,\n",[33,896,897,900,902],{"class":35,"line":110},[33,898,899],{"class":431},"                    value",[33,901,396],{"class":280},[33,903,904],{"class":173},"value,\n",[33,906,907,910,912],{"class":35,"line":115},[33,908,909],{"class":431},"                    reason",[33,911,396],{"class":280},[33,913,914],{"class":358},"\"Extracted value not found verbatim in source — likely fabricated or paraphrased\"\n",[33,916,917],{"class":35,"line":121},[33,918,919],{"class":173},"                )\n",[33,921,922],{"class":35,"line":127},[33,923,71],{"emptyLinePlaceholder":70},[33,925,926,928],{"class":35,"line":133},[33,927,575],{"class":280},[33,929,930],{"class":173}," result\n",[33,932,933],{"class":35,"line":139},[33,934,71],{"emptyLinePlaceholder":70},[33,936,937],{"class":35,"line":145},[33,938,939],{"class":39},"# This catches a REAL, common failure mode: a model paraphrasing or subtly\n",[33,941,942],{"class":35,"line":151},[33,943,944],{"class":39},"# fabricating a \"quote\" that doesn't appear in the source. Far more reliable than\n",[33,946,947],{"class":35,"line":516},[33,948,949],{"class":39},"# asking the model \"is this quote accurate?\" (same model checking its own output).\n",[33,951,952],{"class":35,"line":522},[33,953,71],{"emptyLinePlaceholder":70},[33,955,956],{"class":35,"line":527},[33,957,958],{"class":39},"# GENERAL PRINCIPLE: verification is STRONGEST when checking against ground truth\n",[33,960,961],{"class":35,"line":533},[33,962,963],{"class":39},"# EXTERNAL to the model's generation:\n",[33,965,966],{"class":35,"line":555},[33,967,968],{"class":39},"#   - a database (does this customer ID exist?)\n",[33,970,971],{"class":35,"line":572},[33,972,973],{"class":39},"#   - a calculator (is this arithmetic result correct?)\n",[33,975,976],{"class":35,"line":581},[33,977,978],{"class":39},"#   - the literal source document (does this quote appear verbatim?)\n",[33,980,981],{"class":35,"line":586},[33,982,983],{"class":39},"#   - a schema validator (is this JSON structurally valid?)\n",[33,985,986],{"class":35,"line":607},[33,987,988],{"class":39},"# All of these are cheap, deterministic, and don't depend on the model grading itself.\n",[14,990,992],{"id":991},"multi-model-cross-checking","Multi-Model Cross-Checking",[19,994,996],{"filename":995,"language":22},"multi_prompt_crosscheck.py",[24,997,999],{"className":26,"code":998,"language":22,"meta":28,"style":28},"# Have two differently-prompted calls independently attempt the same task.\n# Agreement = stronger evidence. Disagreement = flag for human judgment.\n\nPROMPT_A = \"\"\"\nWhat is the primary cause of this application crash, based on the attached\nstack trace? State the root cause in one sentence.\n\"\"\"\n\nPROMPT_B = \"\"\"\nA colleague claims the primary cause of this crash is a null pointer\ndereference in the request handler. Review the attached stack trace and\neither confirm or refute that specific claim, citing the exact lines that\nsupport your conclusion.\n\"\"\"\n\nasync def cross_check(stack_trace: str) -> dict:\n    \"\"\"Two independent framings; disagreement triggers human review.\"\"\"\n    result_a = await call_model(PROMPT_A + f\"\\n\\nStack trace:\\n{stack_trace}\")\n    result_b = await call_model(PROMPT_B + f\"\\n\\nStack trace:\\n{stack_trace}\")\n\n    # Compare — if both converge on the same root cause, stronger evidence\n    # If they diverge, flag for human judgment rather than picking arbitrarily\n    if normalize_answer(result_a) == normalize_answer(result_b):\n        return {\"status\": \"agreement\", \"answer\": result_a}\n    else:\n        return {\n            \"status\": \"disagreement\",\n            \"answer_a\": result_a,\n            \"answer_b\": result_b,\n            \"action\": \"escalate_to_human\",\n        }\n\n# CAVEAT: cross-checking catches DIVERGENT errors, not SHARED ones. If both\n# prompts share the same blind spot (both rely on general knowledge about a\n# fact the model is simply wrong about), agreement provides FALSE reassurance.\n# Not a substitute for external grounding (Chapter 12) when the risk is a\n# shared factual gap rather than a reasoning-path fluke.\n",[30,1000,1001,1006,1011,1015,1025,1030,1035,1040,1044,1053,1058,1063,1068,1073,1077,1081,1103,1108,1151,1184,1188,1193,1198,1212,1236,1243,1250,1262,1270,1278,1290,1295,1299,1304,1309,1314,1319],{"__ignoreMap":28},[33,1002,1003],{"class":35,"line":36},[33,1004,1005],{"class":39},"# Have two differently-prompted calls independently attempt the same task.\n",[33,1007,1008],{"class":35,"line":43},[33,1009,1010],{"class":39},"# Agreement = stronger evidence. Disagreement = flag for human judgment.\n",[33,1012,1013],{"class":35,"line":49},[33,1014,71],{"emptyLinePlaceholder":70},[33,1016,1017,1020,1022],{"class":35,"line":55},[33,1018,1019],{"class":316},"PROMPT_A",[33,1021,326],{"class":280},[33,1023,1024],{"class":358}," \"\"\"\n",[33,1026,1027],{"class":35,"line":61},[33,1028,1029],{"class":358},"What is the primary cause of this application crash, based on the attached\n",[33,1031,1032],{"class":35,"line":67},[33,1033,1034],{"class":358},"stack trace? State the root cause in one sentence.\n",[33,1036,1037],{"class":35,"line":74},[33,1038,1039],{"class":358},"\"\"\"\n",[33,1041,1042],{"class":35,"line":80},[33,1043,71],{"emptyLinePlaceholder":70},[33,1045,1046,1049,1051],{"class":35,"line":86},[33,1047,1048],{"class":316},"PROMPT_B",[33,1050,326],{"class":280},[33,1052,1024],{"class":358},[33,1054,1055],{"class":35,"line":92},[33,1056,1057],{"class":358},"A colleague claims the primary cause of this crash is a null pointer\n",[33,1059,1060],{"class":35,"line":98},[33,1061,1062],{"class":358},"dereference in the request handler. Review the attached stack trace and\n",[33,1064,1065],{"class":35,"line":104},[33,1066,1067],{"class":358},"either confirm or refute that specific claim, citing the exact lines that\n",[33,1069,1070],{"class":35,"line":110},[33,1071,1072],{"class":358},"support your conclusion.\n",[33,1074,1075],{"class":35,"line":115},[33,1076,1039],{"class":358},[33,1078,1079],{"class":35,"line":121},[33,1080,71],{"emptyLinePlaceholder":70},[33,1082,1083,1086,1089,1092,1095,1097,1099,1101],{"class":35,"line":127},[33,1084,1085],{"class":280},"async",[33,1087,1088],{"class":280}," def",[33,1090,1091],{"class":309}," cross_check",[33,1093,1094],{"class":173},"(stack_trace: ",[33,1096,317],{"class":316},[33,1098,599],{"class":173},[33,1100,780],{"class":316},[33,1102,604],{"class":173},[33,1104,1105],{"class":35,"line":133},[33,1106,1107],{"class":358},"    \"\"\"Two independent framings; disagreement triggers human review.\"\"\"\n",[33,1109,1110,1113,1115,1118,1121,1123,1126,1129,1132,1135,1138,1141,1144,1147,1149],{"class":35,"line":139},[33,1111,1112],{"class":173},"    result_a ",[33,1114,396],{"class":280},[33,1116,1117],{"class":280}," await",[33,1119,1120],{"class":173}," call_model(",[33,1122,1019],{"class":316},[33,1124,1125],{"class":280}," +",[33,1127,1128],{"class":280}," f",[33,1130,1131],{"class":358},"\"",[33,1133,1134],{"class":316},"\\n\\n",[33,1136,1137],{"class":358},"Stack trace:",[33,1139,1140],{"class":316},"\\n{",[33,1142,1143],{"class":173},"stack_trace",[33,1145,1146],{"class":316},"}",[33,1148,1131],{"class":358},[33,1150,681],{"class":173},[33,1152,1153,1156,1158,1160,1162,1164,1166,1168,1170,1172,1174,1176,1178,1180,1182],{"class":35,"line":145},[33,1154,1155],{"class":173},"    result_b ",[33,1157,396],{"class":280},[33,1159,1117],{"class":280},[33,1161,1120],{"class":173},[33,1163,1048],{"class":316},[33,1165,1125],{"class":280},[33,1167,1128],{"class":280},[33,1169,1131],{"class":358},[33,1171,1134],{"class":316},[33,1173,1137],{"class":358},[33,1175,1140],{"class":316},[33,1177,1143],{"class":173},[33,1179,1146],{"class":316},[33,1181,1131],{"class":358},[33,1183,681],{"class":173},[33,1185,1186],{"class":35,"line":151},[33,1187,71],{"emptyLinePlaceholder":70},[33,1189,1190],{"class":35,"line":516},[33,1191,1192],{"class":39},"    # Compare — if both converge on the same root cause, stronger evidence\n",[33,1194,1195],{"class":35,"line":522},[33,1196,1197],{"class":39},"    # If they diverge, flag for human judgment rather than picking arbitrarily\n",[33,1199,1200,1203,1206,1209],{"class":35,"line":527},[33,1201,1202],{"class":280},"    if",[33,1204,1205],{"class":173}," normalize_answer(result_a) ",[33,1207,1208],{"class":280},"==",[33,1210,1211],{"class":173}," normalize_answer(result_b):\n",[33,1213,1214,1217,1220,1223,1225,1228,1230,1233],{"class":35,"line":533},[33,1215,1216],{"class":280},"        return",[33,1218,1219],{"class":173}," {",[33,1221,1222],{"class":358},"\"status\"",[33,1224,481],{"class":173},[33,1226,1227],{"class":358},"\"agreement\"",[33,1229,348],{"class":173},[33,1231,1232],{"class":358},"\"answer\"",[33,1234,1235],{"class":173},": result_a}\n",[33,1237,1238,1241],{"class":35,"line":555},[33,1239,1240],{"class":280},"    else",[33,1242,604],{"class":173},[33,1244,1245,1247],{"class":35,"line":572},[33,1246,1216],{"class":280},[33,1248,1249],{"class":173}," {\n",[33,1251,1252,1255,1257,1260],{"class":35,"line":581},[33,1253,1254],{"class":358},"            \"status\"",[33,1256,481],{"class":173},[33,1258,1259],{"class":358},"\"disagreement\"",[33,1261,440],{"class":173},[33,1263,1264,1267],{"class":35,"line":586},[33,1265,1266],{"class":358},"            \"answer_a\"",[33,1268,1269],{"class":173},": result_a,\n",[33,1271,1272,1275],{"class":35,"line":607},[33,1273,1274],{"class":358},"            \"answer_b\"",[33,1276,1277],{"class":173},": result_b,\n",[33,1279,1280,1283,1285,1288],{"class":35,"line":613},[33,1281,1282],{"class":358},"            \"action\"",[33,1284,481],{"class":173},[33,1286,1287],{"class":358},"\"escalate_to_human\"",[33,1289,440],{"class":173},[33,1291,1292],{"class":35,"line":619},[33,1293,1294],{"class":173},"        }\n",[33,1296,1297],{"class":35,"line":684},[33,1298,71],{"emptyLinePlaceholder":70},[33,1300,1301],{"class":35,"line":724},[33,1302,1303],{"class":39},"# CAVEAT: cross-checking catches DIVERGENT errors, not SHARED ones. If both\n",[33,1305,1306],{"class":35,"line":729},[33,1307,1308],{"class":39},"# prompts share the same blind spot (both rely on general knowledge about a\n",[33,1310,1311],{"class":35,"line":735},[33,1312,1313],{"class":39},"# fact the model is simply wrong about), agreement provides FALSE reassurance.\n",[33,1315,1316],{"class":35,"line":741},[33,1317,1318],{"class":39},"# Not a substitute for external grounding (Chapter 12) when the risk is a\n",[33,1320,1321],{"class":35,"line":747},[33,1322,1323],{"class":39},"# shared factual gap rather than a reasoning-path fluke.\n",[14,1325,1327],{"id":1326},"where-to-spend-the-verification-budget","Where to Spend the Verification Budget",[19,1329,1331],{"filename":1330,"language":22},"verification_budget.py",[24,1332,1334],{"className":26,"code":1333,"language":22,"meta":28,"style":28},"# Verification costs extra tokens, calls, and latency. Allocate by stakes + measurability:\n\nVERIFICATION_PRIORITIES = {\n    \"early_pipeline_stages\": \"HIGH — an error here propagates through everything downstream\",\n    \"numeric_factual_claims\": \"HIGH — external checks (calculator, source doc) are cheap and available\",\n    \"high_stakes_low_frequency\": \"HIGH — medical\u002Flegal\u002Ffinancial with real consequences\",\n    \"measured_high_error_rate\": \"HIGH — from your eval data (Chapter 9), not intuition\",\n    \"low_stakes_high_frequency\": \"LOW — casual chat; cost of occasional error \u003C cost of verification\",\n    \"creative_generation\": \"LOW — no 'correct' answer to verify against\",\n}\n\n# RULE: spend the verification budget where your DATA says it's needed, not where\n# intuition says it MIGHT be needed. If your eval set shows 2% error on classification\n# but 15% error on extraction, spend your budget on extraction verification.\n",[30,1335,1336,1341,1345,1354,1366,1378,1390,1402,1414,1426,1431,1435,1440,1445],{"__ignoreMap":28},[33,1337,1338],{"class":35,"line":36},[33,1339,1340],{"class":39},"# Verification costs extra tokens, calls, and latency. Allocate by stakes + measurability:\n",[33,1342,1343],{"class":35,"line":43},[33,1344,71],{"emptyLinePlaceholder":70},[33,1346,1347,1350,1352],{"class":35,"line":49},[33,1348,1349],{"class":316},"VERIFICATION_PRIORITIES",[33,1351,326],{"class":280},[33,1353,1249],{"class":173},[33,1355,1356,1359,1361,1364],{"class":35,"line":55},[33,1357,1358],{"class":358},"    \"early_pipeline_stages\"",[33,1360,481],{"class":173},[33,1362,1363],{"class":358},"\"HIGH — an error here propagates through everything downstream\"",[33,1365,440],{"class":173},[33,1367,1368,1371,1373,1376],{"class":35,"line":61},[33,1369,1370],{"class":358},"    \"numeric_factual_claims\"",[33,1372,481],{"class":173},[33,1374,1375],{"class":358},"\"HIGH — external checks (calculator, source doc) are cheap and available\"",[33,1377,440],{"class":173},[33,1379,1380,1383,1385,1388],{"class":35,"line":67},[33,1381,1382],{"class":358},"    \"high_stakes_low_frequency\"",[33,1384,481],{"class":173},[33,1386,1387],{"class":358},"\"HIGH — medical\u002Flegal\u002Ffinancial with real consequences\"",[33,1389,440],{"class":173},[33,1391,1392,1395,1397,1400],{"class":35,"line":74},[33,1393,1394],{"class":358},"    \"measured_high_error_rate\"",[33,1396,481],{"class":173},[33,1398,1399],{"class":358},"\"HIGH — from your eval data (Chapter 9), not intuition\"",[33,1401,440],{"class":173},[33,1403,1404,1407,1409,1412],{"class":35,"line":80},[33,1405,1406],{"class":358},"    \"low_stakes_high_frequency\"",[33,1408,481],{"class":173},[33,1410,1411],{"class":358},"\"LOW — casual chat; cost of occasional error \u003C cost of verification\"",[33,1413,440],{"class":173},[33,1415,1416,1419,1421,1424],{"class":35,"line":86},[33,1417,1418],{"class":358},"    \"creative_generation\"",[33,1420,481],{"class":173},[33,1422,1423],{"class":358},"\"LOW — no 'correct' answer to verify against\"",[33,1425,440],{"class":173},[33,1427,1428],{"class":35,"line":92},[33,1429,1430],{"class":173},"}\n",[33,1432,1433],{"class":35,"line":98},[33,1434,71],{"emptyLinePlaceholder":70},[33,1436,1437],{"class":35,"line":104},[33,1438,1439],{"class":39},"# RULE: spend the verification budget where your DATA says it's needed, not where\n",[33,1441,1442],{"class":35,"line":110},[33,1443,1444],{"class":39},"# intuition says it MIGHT be needed. If your eval set shows 2% error on classification\n",[33,1446,1447],{"class":35,"line":115},[33,1448,1449],{"class":39},"# but 15% error on extraction, spend your budget on extraction verification.\n",[14,1451,1453],{"id":1452},"tips-tricks","💡 Tips & Tricks",[19,1455,1457],{"filename":1456,"language":22},"tips.py",[24,1458,1460],{"className":26,"code":1459,"language":22,"meta":28,"style":28},"# [Idiom] Ask for a different METHOD, not a re-read. \"Verify using a different\n# approach (work backward, check a special case, use a different formula)\" —\n# not \"check your work\" (which too easily becomes a repetition).\n\n# [Idiom] Reserve self-consistency for problems with a CHECKABLE final answer.\n# Majority voting works when answers are comparable (number, category, short fact).\n# It's much harder to apply to open-ended generation (no clean \"majority vote\"\n# over 5 different essays).\n\n# [Debug] Log disagreement rate as a quality signal, not just a routing trigger.\n# A sudden rise in disagreement across samples flags an ambiguous new input\n# pattern or a prompt that's stopped matching current input characteristics.\n\n# [Performance] Cheap external verification beats expensive model-based verification.\n# Checking an extracted quote against source text costs milliseconds of deterministic\n# code. Use it IN PREFERENCE to a second LLM call whenever ground truth is checkable.\n\n# [Performance] Combine self-consistency with decomposition for the highest-stakes\n# SINGLE stage, not the whole pipeline. Running every stage 5x multiplies cost\n# across the entire pipeline — identify the single highest-risk stage and apply\n# heavy verification only there.\n",[30,1461,1462,1467,1472,1477,1481,1486,1491,1496,1501,1505,1510,1515,1520,1524,1529,1534,1539,1543,1548,1553,1558],{"__ignoreMap":28},[33,1463,1464],{"class":35,"line":36},[33,1465,1466],{"class":39},"# [Idiom] Ask for a different METHOD, not a re-read. \"Verify using a different\n",[33,1468,1469],{"class":35,"line":43},[33,1470,1471],{"class":39},"# approach (work backward, check a special case, use a different formula)\" —\n",[33,1473,1474],{"class":35,"line":49},[33,1475,1476],{"class":39},"# not \"check your work\" (which too easily becomes a repetition).\n",[33,1478,1479],{"class":35,"line":55},[33,1480,71],{"emptyLinePlaceholder":70},[33,1482,1483],{"class":35,"line":61},[33,1484,1485],{"class":39},"# [Idiom] Reserve self-consistency for problems with a CHECKABLE final answer.\n",[33,1487,1488],{"class":35,"line":67},[33,1489,1490],{"class":39},"# Majority voting works when answers are comparable (number, category, short fact).\n",[33,1492,1493],{"class":35,"line":74},[33,1494,1495],{"class":39},"# It's much harder to apply to open-ended generation (no clean \"majority vote\"\n",[33,1497,1498],{"class":35,"line":80},[33,1499,1500],{"class":39},"# over 5 different essays).\n",[33,1502,1503],{"class":35,"line":86},[33,1504,71],{"emptyLinePlaceholder":70},[33,1506,1507],{"class":35,"line":92},[33,1508,1509],{"class":39},"# [Debug] Log disagreement rate as a quality signal, not just a routing trigger.\n",[33,1511,1512],{"class":35,"line":98},[33,1513,1514],{"class":39},"# A sudden rise in disagreement across samples flags an ambiguous new input\n",[33,1516,1517],{"class":35,"line":104},[33,1518,1519],{"class":39},"# pattern or a prompt that's stopped matching current input characteristics.\n",[33,1521,1522],{"class":35,"line":110},[33,1523,71],{"emptyLinePlaceholder":70},[33,1525,1526],{"class":35,"line":115},[33,1527,1528],{"class":39},"# [Performance] Cheap external verification beats expensive model-based verification.\n",[33,1530,1531],{"class":35,"line":121},[33,1532,1533],{"class":39},"# Checking an extracted quote against source text costs milliseconds of deterministic\n",[33,1535,1536],{"class":35,"line":127},[33,1537,1538],{"class":39},"# code. Use it IN PREFERENCE to a second LLM call whenever ground truth is checkable.\n",[33,1540,1541],{"class":35,"line":133},[33,1542,71],{"emptyLinePlaceholder":70},[33,1544,1545],{"class":35,"line":139},[33,1546,1547],{"class":39},"# [Performance] Combine self-consistency with decomposition for the highest-stakes\n",[33,1549,1550],{"class":35,"line":145},[33,1551,1552],{"class":39},"# SINGLE stage, not the whole pipeline. Running every stage 5x multiplies cost\n",[33,1554,1555],{"class":35,"line":151},[33,1556,1557],{"class":39},"# across the entire pipeline — identify the single highest-risk stage and apply\n",[33,1559,1560],{"class":35,"line":516},[33,1561,1562],{"class":39},"# heavy verification only there.\n",[14,1564,1566],{"id":1565},"️-edge-cases-gotchas","⚠️ Edge Cases & Gotchas",[19,1568,1570],{"filename":1569,"language":22},"edge_cases.py",[24,1571,1573],{"className":26,"code":1572,"language":22,"meta":28,"style":28},"# [Gotcha] Self-consistency at temperature 0 DOESN'T WORK. If every sample is\n# deterministic, all N samples produce the same output regardless of correctness.\n# Self-consistency REQUIRES actual sampling diversity (nonzero temperature).\n\n# [Gotcha] A confident, articulate wrong answer can WIN a majority vote if the\n# model has a strong, consistent (but mistaken) prior. 5\u002F5 samples agreeing on\n# \"Sydney\" for Australia's capital is NOT proof of correctness — it's a shared\n# misconception. Self-consistency corrects VARIANCE, not BIAS.\n\n# [Gotcha] Verification steps can THEMSELVES introduce errors. A \"double-check\n# this JSON is valid\" pass, run as a model call rather than a deterministic\n# parser, can confidently declare malformed JSON valid or \"fix\" valid JSON into\n# a broken form. Whenever a deterministic check exists, it's strictly more reliable.\n\n# [Gotcha] The cost multiplier is easy to underestimate at scale. 5 samples =\n# 5x token spend AND 5x load on rate limits, EVERY time that code path runs.\n# Reserve explicitly for the subset of requests that need it, not uniformly.\n\n# [Gotcha] Multi-prompt cross-checking can produce two independently wrong answers\n# that AGREE. If both prompts share the same underlying blind spot, agreement\n# provides false reassurance. Cross-checking catches divergent errors, not shared ones.\n",[30,1574,1575,1580,1585,1590,1594,1599,1604,1609,1614,1618,1623,1628,1633,1638,1642,1647,1652,1657,1661,1666,1671],{"__ignoreMap":28},[33,1576,1577],{"class":35,"line":36},[33,1578,1579],{"class":39},"# [Gotcha] Self-consistency at temperature 0 DOESN'T WORK. If every sample is\n",[33,1581,1582],{"class":35,"line":43},[33,1583,1584],{"class":39},"# deterministic, all N samples produce the same output regardless of correctness.\n",[33,1586,1587],{"class":35,"line":49},[33,1588,1589],{"class":39},"# Self-consistency REQUIRES actual sampling diversity (nonzero temperature).\n",[33,1591,1592],{"class":35,"line":55},[33,1593,71],{"emptyLinePlaceholder":70},[33,1595,1596],{"class":35,"line":61},[33,1597,1598],{"class":39},"# [Gotcha] A confident, articulate wrong answer can WIN a majority vote if the\n",[33,1600,1601],{"class":35,"line":67},[33,1602,1603],{"class":39},"# model has a strong, consistent (but mistaken) prior. 5\u002F5 samples agreeing on\n",[33,1605,1606],{"class":35,"line":74},[33,1607,1608],{"class":39},"# \"Sydney\" for Australia's capital is NOT proof of correctness — it's a shared\n",[33,1610,1611],{"class":35,"line":80},[33,1612,1613],{"class":39},"# misconception. Self-consistency corrects VARIANCE, not BIAS.\n",[33,1615,1616],{"class":35,"line":86},[33,1617,71],{"emptyLinePlaceholder":70},[33,1619,1620],{"class":35,"line":92},[33,1621,1622],{"class":39},"# [Gotcha] Verification steps can THEMSELVES introduce errors. A \"double-check\n",[33,1624,1625],{"class":35,"line":98},[33,1626,1627],{"class":39},"# this JSON is valid\" pass, run as a model call rather than a deterministic\n",[33,1629,1630],{"class":35,"line":104},[33,1631,1632],{"class":39},"# parser, can confidently declare malformed JSON valid or \"fix\" valid JSON into\n",[33,1634,1635],{"class":35,"line":110},[33,1636,1637],{"class":39},"# a broken form. Whenever a deterministic check exists, it's strictly more reliable.\n",[33,1639,1640],{"class":35,"line":115},[33,1641,71],{"emptyLinePlaceholder":70},[33,1643,1644],{"class":35,"line":121},[33,1645,1646],{"class":39},"# [Gotcha] The cost multiplier is easy to underestimate at scale. 5 samples =\n",[33,1648,1649],{"class":35,"line":127},[33,1650,1651],{"class":39},"# 5x token spend AND 5x load on rate limits, EVERY time that code path runs.\n",[33,1653,1654],{"class":35,"line":133},[33,1655,1656],{"class":39},"# Reserve explicitly for the subset of requests that need it, not uniformly.\n",[33,1658,1659],{"class":35,"line":139},[33,1660,71],{"emptyLinePlaceholder":70},[33,1662,1663],{"class":35,"line":145},[33,1664,1665],{"class":39},"# [Gotcha] Multi-prompt cross-checking can produce two independently wrong answers\n",[33,1667,1668],{"class":35,"line":151},[33,1669,1670],{"class":39},"# that AGREE. If both prompts share the same underlying blind spot, agreement\n",[33,1672,1673],{"class":35,"line":516},[33,1674,1675],{"class":39},"# provides false reassurance. Cross-checking catches divergent errors, not shared ones.\n",[14,1677,1679],{"id":1678},"spot-the-bug","🧠 Spot the Bug",[1681,1682,1683],"p",{},"A medical triage assistant adds: \"You just classified this patient's symptoms as LOW urgency. Before finalizing, double-check: are you confident this is correct?\" The check almost never changes the classification — the model always responds \"Yes, I'm confident.\" The team concludes it's working well. What's the flaw?",[1685,1686,1687,1691,1699,1702,1725],"details",{},[1688,1689,1690],"summary",{},"Answer",[1681,1692,1693,1694,1698],{},"A verification step that ",[1695,1696,1697],"em",{},"always confirms"," the original answer provides zero information — it can't distinguish \"the classification was actually correct\" from \"the verification step is a rubber stamp that never meaningfully re-examines anything.\" This is the naive self-verification failure: asking the same model, with the same reasoning, whether it's \"confident\" introduces no new derivation or external signal that could catch an error.",[1681,1700,1701],{},"For a high-stakes domain like medical triage, a rigorous design would use:",[1703,1704,1705,1713,1719],"ol",{},[1706,1707,1708,1712],"li",{},[1709,1710,1711],"strong",{},"Structured, method-different verification"," — re-derive the urgency from symptoms against an explicit checklist of red-flag symptoms, independent of the first pass's reasoning.",[1706,1714,1715,1718],{},[1709,1716,1717],{},"An external, non-LLM safety net"," — a hard rule that any symptom matching a predefined red-flag list is escalated to at least MEDIUM urgency regardless of the model's classification.",[1706,1720,1721,1724],{},[1709,1722,1723],{},"Human clinical review"," as the actual safeguard for anything the system is uncertain about.",[1681,1726,1727],{},"A verification step that virtually always confirms the original answer is not evidence the answer was right — it's a sign the verification step isn't introducing any new reasoning path or external check.",[14,1729,1731],{"id":1730},"key-takeaways","Key Takeaways",[19,1733,1735],{"filename":1734,"language":22},"key_takeaways.py",[24,1736,1738],{"className":26,"code":1737,"language":22,"meta":28,"style":28},"\"\"\"\nSelf-consistency & verification — catching errors beyond single-run.\n\"\"\"\n\n# 1. Naive self-verification (\"are you sure?\") catches surface-visible errors\n#    (arithmetic slips, malformed structure) but NOT errors rooted in what the\n#    model actually believes. Same reasoning re-examined → same conclusion.\n\n# 2. Structured verification forcing a DIFFERENT METHOD is meaningfully stronger:\n#    a different derivation arriving at the same answer is real corroboration.\n\n# 3. Self-consistency (N samples at nonzero temp, majority vote) corrects VARIANCE\n#    but NOT BIAS. 5\u002F5 agreeing on a shared misconception ≠ proof of correctness.\n\n# 4. Verification against EXTERNAL ground truth (source docs, calculators, schema\n#    validators, databases) is more reliable than ANY form of self-checking.\n#    Prefer external checks whenever ground truth is mechanically accessible.\n\n# 5. Verification budget is finite — allocate by measured error rates and stakes.\n#    Prioritize: early pipeline stages, checkable factual claims, high-stakes\n#    low-frequency decisions. Don't apply uniformly — cost compounds at scale.\n",[30,1739,1740,1744,1749,1753,1757,1762,1767,1772,1776,1781,1786,1790,1795,1800,1804,1809,1814,1819,1823,1828,1833],{"__ignoreMap":28},[33,1741,1742],{"class":35,"line":36},[33,1743,1039],{"class":358},[33,1745,1746],{"class":35,"line":43},[33,1747,1748],{"class":358},"Self-consistency & verification — catching errors beyond single-run.\n",[33,1750,1751],{"class":35,"line":49},[33,1752,1039],{"class":358},[33,1754,1755],{"class":35,"line":55},[33,1756,71],{"emptyLinePlaceholder":70},[33,1758,1759],{"class":35,"line":61},[33,1760,1761],{"class":39},"# 1. Naive self-verification (\"are you sure?\") catches surface-visible errors\n",[33,1763,1764],{"class":35,"line":67},[33,1765,1766],{"class":39},"#    (arithmetic slips, malformed structure) but NOT errors rooted in what the\n",[33,1768,1769],{"class":35,"line":74},[33,1770,1771],{"class":39},"#    model actually believes. Same reasoning re-examined → same conclusion.\n",[33,1773,1774],{"class":35,"line":80},[33,1775,71],{"emptyLinePlaceholder":70},[33,1777,1778],{"class":35,"line":86},[33,1779,1780],{"class":39},"# 2. Structured verification forcing a DIFFERENT METHOD is meaningfully stronger:\n",[33,1782,1783],{"class":35,"line":92},[33,1784,1785],{"class":39},"#    a different derivation arriving at the same answer is real corroboration.\n",[33,1787,1788],{"class":35,"line":98},[33,1789,71],{"emptyLinePlaceholder":70},[33,1791,1792],{"class":35,"line":104},[33,1793,1794],{"class":39},"# 3. Self-consistency (N samples at nonzero temp, majority vote) corrects VARIANCE\n",[33,1796,1797],{"class":35,"line":110},[33,1798,1799],{"class":39},"#    but NOT BIAS. 5\u002F5 agreeing on a shared misconception ≠ proof of correctness.\n",[33,1801,1802],{"class":35,"line":115},[33,1803,71],{"emptyLinePlaceholder":70},[33,1805,1806],{"class":35,"line":121},[33,1807,1808],{"class":39},"# 4. Verification against EXTERNAL ground truth (source docs, calculators, schema\n",[33,1810,1811],{"class":35,"line":127},[33,1812,1813],{"class":39},"#    validators, databases) is more reliable than ANY form of self-checking.\n",[33,1815,1816],{"class":35,"line":133},[33,1817,1818],{"class":39},"#    Prefer external checks whenever ground truth is mechanically accessible.\n",[33,1820,1821],{"class":35,"line":139},[33,1822,71],{"emptyLinePlaceholder":70},[33,1824,1825],{"class":35,"line":145},[33,1826,1827],{"class":39},"# 5. Verification budget is finite — allocate by measured error rates and stakes.\n",[33,1829,1830],{"class":35,"line":151},[33,1831,1832],{"class":39},"#    Prioritize: early pipeline stages, checkable factual claims, high-stakes\n",[33,1834,1835],{"class":35,"line":516},[33,1836,1837],{"class":39},"#    low-frequency decisions. Don't apply uniformly — cost compounds at scale.\n",[1839,1840,1841],"style",{},"html pre.shiki code .sdCPZ, html code.shiki .sdCPZ{--shiki-default:#6A737D;--shiki-github-dark:#6A737D}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .github-dark .shiki span {color: var(--shiki-github-dark);background: var(--shiki-github-dark-bg);font-style: var(--shiki-github-dark-font-style);font-weight: var(--shiki-github-dark-font-weight);text-decoration: var(--shiki-github-dark-text-decoration);}html.github-dark .shiki span {color: var(--shiki-github-dark);background: var(--shiki-github-dark-bg);font-style: var(--shiki-github-dark-font-style);font-weight: var(--shiki-github-dark-font-weight);text-decoration: var(--shiki-github-dark-text-decoration);}html pre.shiki code .ssxIu, html code.shiki .ssxIu{--shiki-default:#24292E;--shiki-github-dark:#E1E4E8}html pre.shiki code .svdQ7, html code.shiki .svdQ7{--shiki-default:#D73A49;--shiki-github-dark:#F97583}html pre.shiki code .sIsaT, html code.shiki .sIsaT{--shiki-default:#6F42C1;--shiki-github-dark:#B392F0}html pre.shiki code .snvgF, html code.shiki .snvgF{--shiki-default:#005CC5;--shiki-github-dark:#79B8FF}html pre.shiki code .sJ6F3, html code.shiki .sJ6F3{--shiki-default:#032F62;--shiki-github-dark:#9ECBFF}html pre.shiki code .sCrzJ, html code.shiki .sCrzJ{--shiki-default:#E36209;--shiki-github-dark:#FFAB70}html pre.shiki code .svAP2, html code.shiki .svAP2{--shiki-default:#032F62;--shiki-github-dark:#DBEDFF}html pre.shiki code .snRuI, html code.shiki .snRuI{--shiki-default:#22863A;--shiki-default-font-weight:bold;--shiki-github-dark:#85E89D;--shiki-github-dark-font-weight:bold}",{"title":28,"searchDepth":43,"depth":43,"links":1843},[1844,1845,1846,1847,1848,1849,1850,1851,1852,1853],{"id":16,"depth":43,"text":17},{"id":157,"depth":43,"text":158},{"id":266,"depth":43,"text":267},{"id":753,"depth":43,"text":754},{"id":991,"depth":43,"text":992},{"id":1326,"depth":43,"text":1327},{"id":1452,"depth":43,"text":1453},{"id":1565,"depth":43,"text":1566},{"id":1678,"depth":43,"text":1679},{"id":1730,"depth":43,"text":1731},"Sampling multiple independent reasoning paths and voting — naive self-verification vs structured verification, self-consistency, external ground-truth checks, and multi-model cross-checking. Code-first reference for mid-to-senior engineers.","md",{},"\u002Fprompt-engineering\u002F11-self-consistency-and-verification",{"title":5,"description":1854},"prompt-engineering\u002F11-self-consistency-and-verification","YBdNDHjiYhWbzQ8Bmp-ScbBZjjqbcq5hHVtQOjGnPkY",1789924650965]