AI 一本正经地编数据,这是用它做投研最大的坑
说个很多人踩过的雷:问 AI 某个币过去一年的最大回撤,它张口就给你一个精确到小数点两位的数字,看着专业极了——但那是编的。这篇讲怎么防。
**I. First, acknowledge a fact**
The essence of a large language model is word chaining, not a database. It has no instinct to say “I can’t find that.” When it lacks data, it fills in with statistically most plausible numbers—and it does so with extra confidence. Real-time or precise data like price, market cap, and trading volume are the worst-hit areas. If you treat it like a market terminal, it will perform “as if” it were one for you.
**II. Three lines of defense**
1️⃣ Get the role right first. Start by saying: “You don’t have real-time quotes. For anything you’re not sure about, say you don’t know—don’t estimate and output it as fact.” This alone can block a large share of hallucinations.
2️⃣ Make it distinguish “facts” from “inferences.” Add to your prompt: “In your answer, split all numbers into two categories—those that come from the materials I provide, and those you infer yourself. Label them separately.” Inferences aren’t forbidden, but you need to know which parts are inferences.
3️⃣ Verify key numbers. For any numbers that need to drive decisions—price, fees, unlocked amounts—check against the exchange pages or a data site. The AI’s job is to ask the right questions thoroughly; your job is the final acceptance test. It’s the same idea as writing code with AI: the more polished the output looks, the more you need to verify.
**III. The right division of labor**
AI’s real strength isn’t reporting numbers—it’s processing the numbers you give it: parse CSV, compute distributions, find abnormal trades, compress a 50-page English research report into three paragraphs of summary—the raw material is yours, and it does the processing. Any step where it needs to fabricate data out of thin air should be treated with default suspicion.
There’s also a simple “low-tech” method that works well: ask the same number twice, with a few other conversations in between. If the two answers differ, you can basically conclude it’s making things up—numbers it truly “remembers” won’t drift. Ten seconds of extra cost can stop the most low-effort kind of hallucination.
The smarter the tool, the more valuable your verification is. These days, the scarcest people aren’t those who can use AI—they’re those who use AI and then still cross-check.
Have you ever been caught by a data pit created by AI? Share a case in the comments.
$BTC #加密
说个很多人踩过的雷:问 AI 某个币过去一年的最大回撤,它张口就给你一个精确到小数点两位的数字,看着专业极了——但那是编的。这篇讲怎么防。
**I. First, acknowledge a fact**
The essence of a large language model is word chaining, not a database. It has no instinct to say “I can’t find that.” When it lacks data, it fills in with statistically most plausible numbers—and it does so with extra confidence. Real-time or precise data like price, market cap, and trading volume are the worst-hit areas. If you treat it like a market terminal, it will perform “as if” it were one for you.
**II. Three lines of defense**
1️⃣ Get the role right first. Start by saying: “You don’t have real-time quotes. For anything you’re not sure about, say you don’t know—don’t estimate and output it as fact.” This alone can block a large share of hallucinations.
2️⃣ Make it distinguish “facts” from “inferences.” Add to your prompt: “In your answer, split all numbers into two categories—those that come from the materials I provide, and those you infer yourself. Label them separately.” Inferences aren’t forbidden, but you need to know which parts are inferences.
3️⃣ Verify key numbers. For any numbers that need to drive decisions—price, fees, unlocked amounts—check against the exchange pages or a data site. The AI’s job is to ask the right questions thoroughly; your job is the final acceptance test. It’s the same idea as writing code with AI: the more polished the output looks, the more you need to verify.
**III. The right division of labor**
AI’s real strength isn’t reporting numbers—it’s processing the numbers you give it: parse CSV, compute distributions, find abnormal trades, compress a 50-page English research report into three paragraphs of summary—the raw material is yours, and it does the processing. Any step where it needs to fabricate data out of thin air should be treated with default suspicion.
There’s also a simple “low-tech” method that works well: ask the same number twice, with a few other conversations in between. If the two answers differ, you can basically conclude it’s making things up—numbers it truly “remembers” won’t drift. Ten seconds of extra cost can stop the most low-effort kind of hallucination.
The smarter the tool, the more valuable your verification is. These days, the scarcest people aren’t those who can use AI—they’re those who use AI and then still cross-check.
Have you ever been caught by a data pit created by AI? Share a case in the comments.
$BTC #加密