Skip to main content
All modules

VerifierResearchresearchacademic writingjournalism

Claim Verifier

Review a claim and its citations without reducing evidence to a truth badge: source identity and status, exact claim support, citation coverage, and consistency with the wider record.

claim-to-check.mdMarkdown
# Claim to check

Paste a passage here, or select text in any document and choose **Verify**.

Paste a claim into claim-to-check.md or select text in any document and use Verify. Claim Verifier will give you one short Contradicted or Not contradicted result, with the full evidence review available in its thread. It will not change your document unless you ask.

What’s inside

  • Skills
  • skills/claim-verification/SKILL.mdclaim-verification · 5.8 KBUse when the user asks to verify, fact-check, validate, audit, or review a claim or citation. Checks bibliographic identity and status, exact claim-to-source support, citation coverage, and consistency with workspace and external evidence while making uncertainty explicit.
    ---
    name: claim-verification
    description: >
      Use when the user asks to verify, fact-check, validate, audit, or review a
      claim or citation. Checks bibliographic identity and status, exact
      claim-to-source support, citation coverage, and consistency with workspace
      and external evidence while making uncertainty explicit.
    ---
    
    # Claim Verifier
    
    Verification is an evidence review, not a truth badge. Keep these questions
    separate throughout the investigation:
    
    1. **Reference integrity:** Does the cited object exist, and do its author,
       title, date, DOI, version, and publication status match the document?
    2. **Claim support:** Does an exact passage in that source support this exact
       claim, with the same population, conditions, magnitude, and certainty?
    3. **Citation coverage:** Are checkable statements cited, and are citations
       placed on the claims they support rather than merely nearby?
    4. **External consistency:** What do relevant primary or authoritative sources
       say beyond the cited source?
    
    ## Workflow
    
    ### 1. Atomize and scope
    
    - Break the selection into independently checkable claims. Separate facts from
      forecasts, opinions, definitions, and recommendations.
    - Read enough surrounding workspace context to resolve pronouns, qualifiers,
      citation markers, dates, and units. Do not silently broaden the claim.
    - State what kind of check is possible. A private, inaccessible, or uncited
      source limits the result; it does not make the claim false.
    
    ### 2. Inspect workspace evidence first
    
    - Find the cited entry and any attached source using `Glob`, `Grep`, and
      `Read`. Check in-text marker ↔ bibliography consistency before searching.
    - For quotations, confirm the wording and locator. For paraphrases, compare
      meaning rather than keyword overlap.
    - Record the shortest decisive evidence passage, page/section when available,
      source version, and access date. Read the original source; snippets and
      search-result summaries are discovery aids only.
    
    ### 3. Check source identity and status
    
    - For Markdown or LaTeX citation structure, optionally run
      `python3 <this-skill-folder>/scripts/audit_citations.py --document '<path>'
      --format json`, resolving `<this-skill-folder>` from the `SKILL.md` path you
      read. The script is offline: it checks citation keys,
      bibliography entries, duplicate identifiers, DOI strings, and URLs, but
      never retrieves evidence or decides whether a source supports the claim.
    - Use Crossref first for publication metadata, publisher-deposited updates, and
      Retraction Watch coverage; DataCite for datasets/software and other DataCite
      objects; OpenAlex for discovery/citation graph context; Unpaywall only to
      locate a lawful open copy.
    - Registry metadata may be incomplete. A clean lookup means no issue was found
      in the records checked, not that the source is valid in every respect.
    - Read `references/providers.md` when choosing or interpreting an external
      provider, especially optional Scite, Elicit, Consensus, or Fact Check lookup.
    
    ### 4. Test exact support
    
    - Compare each atomic claim to exact source passages. Check entity/population,
      intervention or condition, outcome, time period, direction, magnitude, and
      uncertainty. Watch for causal language derived from correlational evidence.
    - Prefer primary sources for factual findings and original attribution;
      authoritative official sources for laws, statistics, and current policy;
      reviews for context, not as automatic replacements for originals.
    - Search externally with the agent's web search/fetch tools when workspace
      evidence is absent, ambiguous, stale, or contested. Use multiple independent
      search formulations and follow results to primary sources. Do not use Bash as
      a substitute for web access. If recency matters, verify dates and status.
    - Google Fact Check results only show that a publisher previously reviewed a
      similar public claim. No result means “no prior review found,” never “true.”
    
    ### 5. Report, then stop
    
    Begin with this exact two-part contract so the app can show one small inline
    result while retaining the full review in its durable thread:
    
    ```markdown
    ## Claim Verifier result
    Contradicted — <one short reason>
    
    ## Details
    <complete evidence review>
    ```
    
    The result line must be exactly one line, keep its reason to at most 12 words,
    and begin with either
    `Contradicted —` or `Not contradicted —`. Use **Contradicted** only when reliable
    checked evidence directly conflicts with at least one material checkable claim
    in the selected passage. Otherwise use **Not contradicted**, including when the
    evidence is incomplete or inaccessible; state that limitation in the short
    reason. “Not contradicted” does not mean supported or verified true. Do not use
    any other overall verdict in the result line.
    
    Under `## Details`, use one finding per atomic claim:
    
    - **Supported** — the checked evidence directly supports the material claim.
    - **Partially supported** — a narrower or qualified form is supported.
    - **Contradicted** — reliable checked evidence directly conflicts with it.
    - **Insufficient evidence** — available evidence cannot resolve it.
    - **Source inaccessible** — the necessary source or passage could not be read.
    - **Needs human review** — interpretation, domain expertise, or conflicting
      high-quality evidence prevents a responsible automated conclusion.
    
    For each finding include:
    
    - the atomic claim and verdict;
    - reference integrity/status separately from support;
    - a brief exact evidence excerpt or precise paraphrase with link and locator;
    - reasoning about material scope differences;
    - sources/providers checked, access date, and unresolved limitations.
    
    Finish with a compact citation-consistency checklist for the selected passage.
    Say “No issue found in the sources checked” when appropriate, never “verified
    true.” Do not modify the document unless the user explicitly asks for edits.
    
  • Starter files
  • claim-to-check.mdopens first · 95 B
    # Claim to check
    
    Paste a passage here, or select text in any document and choose **Verify**.
    
    
  • skills/claim-verification/references/providers.md2.7 KB
    # Evidence provider guide
    
    No provider covers every verification layer. Use the smallest relevant set and
    retain the exact provider response or source passage in the audit trail.
    
    ## Baseline public services
    
    | Service | Use it for | Do not infer |
    | --- | --- | --- |
    | [Crossref](https://www.crossref.org/documentation/retrieve-metadata/rest-api/) | DOI identity, publisher-deposited metadata, relations and updates | Full-text support or source quality |
    | [Retraction Watch via Crossref](https://www.crossref.org/documentation/retrieve-metadata/retraction-watch/) and [Crossmark](https://www.crossref.org/services/crossmark/) | Retractions and participating publishers' current-version notices | Complete correction/expression-of-concern coverage |
    | [DataCite](https://support.datacite.org/docs/rest-api) | DOI metadata for datasets, software, and non-Crossref objects | Peer review or claim support |
    | [OpenAlex](https://developers.openalex.org/) | Work discovery, citation graph context, and some open full text | A definitive publication-status verdict |
    | [Unpaywall](https://unpaywall.org/products/api) | A lawful open-access location and visible manuscript version | Quality, correctness, or support |
    | [Google Fact Check Tools](https://developers.google.com/fact-check/tools/api/reference/rest/v1alpha1/claims/search) | Prior reviews of similar public claims | Truth from a match, or truth from no match |
    
    Use the agent's web search/fetch tools to open registry records, publisher
    pages, full text, and primary evidence. Do not use Bash, Python HTTP, `curl`, or
    `wget` as a substitute for unavailable web tools. The bundled
    `scripts/audit_citations.py` is intentionally offline and structural only.
    
    ## Optional connected services
    
    - **Scite:** later-paper citation contexts and supporting/contrasting/mentioning
      signals. These are useful leads, not proof that the user's cited passage
      supports their sentence.
    - **Elicit:** evidence extraction with an exact supporting quote. Re-read the
      original and preserve the quote/locator; its AI output remains reviewable.
    - **Consensus:** literature discovery and evidence landscapes. Use it to find
      studies, then verify the original sources.
    - **Zotero:** library identity, linked citations, locators, bibliography refresh,
      and retraction warnings. It does not judge claim support.
    
    Only use an optional service when its connector or credentials are already
    available in the workspace. Never request secrets in chat or invent access.
    
    ## Retrieval order
    
    1. Attached workspace source and cited passage.
    2. DOI registry and publisher/current-version page.
    3. Lawful full text or authoritative primary source.
    4. Independent primary/context sources.
    5. Prior fact checks or citation-network signals as supplemental context.
    
  • skills/claim-verification/scripts/audit_citations.py8.0 KB
    #!/usr/bin/env python3
    """Offline Markdown/LaTeX citation consistency audit; no network access."""
    
    from __future__ import annotations
    
    import argparse
    import json
    import re
    import sys
    from pathlib import Path
    from typing import Any
    
    CITE_RE = re.compile(
        r"\\(?:cite|citep|citet|citealp|citeauthor|citeyear|autocite|parencite|textcite|footcite)\w*"
        r"\s*(?:\[[^\]]*\]\s*){0,2}\{([^}]+)\}"
    )
    PANDOC_CITE_RE = re.compile(r"(?<![\w@])@([A-Za-z0-9_:.+/-]+)")
    DOI_RE = re.compile(r"(?i)\b10\.\d{4,9}/[-._;()/:A-Z0-9]+")
    URL_RE = re.compile(r"https?://[^\s<>\]\[{}\"']+")
    
    
    def clean_doi(value: str) -> str:
        value = re.sub(r"(?i)^(?:https?://(?:dx\.)?doi\.org/|doi:\s*)", "", value.strip())
        return value.rstrip(".,;:)]}").lower()
    
    
    def line_number(text: str, offset: int) -> int:
        return text.count("\n", 0, offset) + 1
    
    
    def extract_document(text: str, start: int | None, end: int | None) -> dict[str, Any]:
        citations: list[dict[str, Any]] = []
        seen_occurrences: set[tuple[str, int]] = set()
        for match in CITE_RE.finditer(text):
            line = line_number(text, match.start())
            if start is not None and not start <= line <= end:  # type: ignore[operator]
                continue
            for raw_key in match.group(1).split(","):
                key = raw_key.strip()
                if key and (key, line) not in seen_occurrences:
                    citations.append({"key": key, "line": line})
                    seen_occurrences.add((key, line))
        for match in PANDOC_CITE_RE.finditer(text):
            line = line_number(text, match.start())
            if start is not None and not start <= line <= end:  # type: ignore[operator]
                continue
            key = match.group(1).rstrip(".,;:")
            if (key, line) not in seen_occurrences:
                citations.append({"key": key, "line": line})
                seen_occurrences.add((key, line))
    
        bib_paths: list[str] = []
        for match in re.finditer(r"\\bibliography\{([^}]+)\}", text):
            bib_paths.extend(
                value if value.lower().endswith(".bib") else f"{value}.bib"
                for value in (part.strip() for part in match.group(1).split(","))
                if value
            )
        bib_paths.extend(
            match.group(1).strip()
            for match in re.finditer(r"\\addbibresource(?:\[[^\]]*\])?\{([^}]+)\}", text)
        )
        yaml_bib = re.search(r"(?ms)^---\s*$.*?^bibliography:\s*([^\n]+).*?^---\s*$", text)
        if yaml_bib:
            value = yaml_bib.group(1).strip().strip("[]")
            bib_paths.extend(part.strip().strip("'\"") for part in value.split(",") if part.strip())
    
        return {
            "citation_occurrences": citations,
            "bibliography_paths": list(dict.fromkeys(bib_paths)),
            "dois": sorted({clean_doi(match.group(0)) for match in DOI_RE.finditer(text)}),
            "urls": sorted({match.group(0).rstrip(".,;:)") for match in URL_RE.finditer(text)}),
        }
    
    
    def bib_entries(text: str, source: str) -> tuple[list[dict[str, Any]], list[str]]:
        entries: list[dict[str, Any]] = []
        warnings: list[str] = []
        cursor = 0
        header = re.compile(r"@([A-Za-z]+)\s*([({])\s*([^,\s]+)\s*,")
        while match := header.search(text, cursor):
            entry_type, opener, key = match.group(1).lower(), match.group(2), match.group(3)
            closer = "}" if opener == "{" else ")"
            depth = 1
            quote = False
            escaped = False
            index = match.end()
            while index < len(text) and depth:
                char = text[index]
                if escaped:
                    escaped = False
                elif char == "\\":
                    escaped = True
                elif char == '"':
                    quote = not quote
                elif not quote and char == opener:
                    depth += 1
                elif not quote and char == closer:
                    depth -= 1
                index += 1
            if depth:
                warnings.append(f"{source}:{line_number(text, match.start())}: unterminated @{entry_type}{{{key}")
                break
            body = text[match.end() : index - 1]
            fields: dict[str, str] = {}
            field_re = re.compile(r"(?im)^\s*([A-Za-z][\w-]*)\s*=\s*")
            matches = list(field_re.finditer(body))
            for pos, field_match in enumerate(matches):
                value = body[field_match.end() : matches[pos + 1].start() if pos + 1 < len(matches) else len(body)]
                value = value.strip().rstrip(",").strip()
                if len(value) >= 2 and ((value[0], value[-1]) in {("{", "}"), ('"', '"')}):
                    value = value[1:-1].strip()
                fields[field_match.group(1).lower()] = re.sub(r"\s+", " ", value)
            entries.append(
                {
                    "key": key,
                    "type": entry_type,
                    "title": fields.get("title"),
                    "author": fields.get("author"),
                    "year": fields.get("year"),
                    "doi": clean_doi(fields["doi"]) if fields.get("doi") else None,
                    "url": fields.get("url"),
                    "source": source,
                }
            )
            cursor = index
        return entries, warnings
    
    
    def main() -> int:
        parser = argparse.ArgumentParser(description=__doc__)
        parser.add_argument("--document", required=True)
        parser.add_argument("--bib", action="append", default=[])
        parser.add_argument("--start-line", type=int)
        parser.add_argument("--end-line", type=int)
        parser.add_argument("--format", choices=("json", "text"), default="json")
        parser.add_argument("--strict", action="store_true")
        args = parser.parse_args()
        if (args.start_line is None) != (args.end_line is None):
            parser.error("--start-line and --end-line must be supplied together")
        if args.start_line is not None and (args.start_line < 1 or args.end_line < args.start_line):
            parser.error("invalid line range")
    
        document = Path(args.document)
        try:
            text = document.read_text(encoding="utf-8")
        except OSError as error:
            parser.error(str(error))
        extracted = extract_document(text, args.start_line, args.end_line)
        requested_bibs = list(dict.fromkeys([*args.bib, *extracted.pop("bibliography_paths")]))
        entries: list[dict[str, Any]] = []
        warnings: list[str] = []
        for raw_path in requested_bibs:
            path = Path(raw_path)
            if not path.is_absolute():
                path = document.parent / path
            try:
                parsed, parse_warnings = bib_entries(path.read_text(encoding="utf-8"), str(path))
                entries.extend(parsed)
                warnings.extend(parse_warnings)
            except OSError as error:
                warnings.append(f"{path}: {error}")
    
        key_counts: dict[str, int] = {}
        doi_counts: dict[str, int] = {}
        for entry in entries:
            key_counts[entry["key"]] = key_counts.get(entry["key"], 0) + 1
            if entry["doi"]:
                doi_counts[entry["doi"]] = doi_counts.get(entry["doi"], 0) + 1
        known_keys = set(key_counts)
        cited_keys = {item["key"] for item in extracted["citation_occurrences"]}
        report = {
            "schema_version": 1,
            "document": str(document),
            **extracted,
            "bibliography_entries": entries,
            "missing_keys": sorted(cited_keys - known_keys) if requested_bibs else [],
            "duplicate_keys": sorted(key for key, count in key_counts.items() if count > 1),
            "duplicate_dois": sorted(doi for doi, count in doi_counts.items() if count > 1),
            "parse_warnings": warnings,
            "scope": "structural only; no URLs fetched and no claim-support verdicts assigned",
        }
        findings = bool(report["missing_keys"] or report["duplicate_keys"] or report["duplicate_dois"] or warnings)
        if args.format == "json":
            json.dump(report, sys.stdout, indent=2, ensure_ascii=False)
            sys.stdout.write("\n")
        else:
            print(f"{len(report['citation_occurrences'])} citations; {len(entries)} bibliography entries")
            for key in ("missing_keys", "duplicate_keys", "duplicate_dois", "parse_warnings"):
                if report[key]:
                    print(f"{key}: {', '.join(report[key])}")
        return 1 if args.strict and findings else 0
    
    
    if __name__ == "__main__":
        raise SystemExit(main())