What do you use as “ground truth”? The page says “independent sources”, and I’m sure there’s too many to list, but my question is how are they vetted as being truthful and how are two sources with opposite viewpoints reconciled?
Something I've found surprisingly effective is telling ChatGPT to "use credible sources" - you can then watch its thinking trace and see it do things like ruling out random blogs, considering media publications with a good reputation for fact checking, and double-checking information that seems unlikely.
There is no whitelist source. It can't rate sources for truthfulness.
When sources conflicts - skill drops verdict to misleading or unverifiable and both are linked. Can't pick a winner at the moment.
That's probably the weakest part and still a judgement call for a model.
At the moment - it's model's judgement upon search results for every claim from provided content. If search results returned inaccurate data, or were fabricated - it might affect a verdict for that claim, or highlight that it's unverifiable or contradictory.
But there is a room for improvement, what do you suggest? Have a Judge agent which will check results?
That's fair. Install is a one time real friction and extension would be much simpler.
But, there is a reason why it's not an extension. At least for now. Whole product isn't a UI or service, it's a set of skills for your AI agent of choice, which gets content from URL or whatever shared as a source, splits it into claims and does extensive web searches for every claim to compare it with statement from provided content.
On local models constraint isn't a model itself - but search. Model don't judge from it's trained memory, so even local model will need a backend for search, otherwise it can't provide a verdict.
Washington Post gave up tracking Trump lies last term in 2021 because it became impossible by human hands with 21+ per day and over 30,000 in their database
but with "AI" now it's possible not only to do non-stop but in REALTIME
you could even just restrict the source of the check to the paper's own reporting the past fifty years
Hmm, smells AI generated. Why should an LLM (this is from a SKILL.md) care about $LINES?
But there is a room for improvement, what do you suggest? Have a Judge agent which will check results?
But, there is a reason why it's not an extension. At least for now. Whole product isn't a UI or service, it's a set of skills for your AI agent of choice, which gets content from URL or whatever shared as a source, splits it into claims and does extensive web searches for every claim to compare it with statement from provided content.
On local models constraint isn't a model itself - but search. Model don't judge from it's trained memory, so even local model will need a backend for search, otherwise it can't provide a verdict.
HN’s community has built-in bullshit detectors.
but with "AI" now it's possible not only to do non-stop but in REALTIME
you could even just restrict the source of the check to the paper's own reporting the past fifty years
* https://www.washingtonpost.com/graphics/politics/trump-claim...
I'd like to see that backfilled, all the way back to the "long form birth certificate" (remember that horror show)