Data Integrity

Our Methodology & Integrity Rules

The only serious risk we see is publishing unreliable data. Here is how we prevent it.

Methodology v1.4— Last updated July 2026

When methodology changes, people must know. Every version is documented.

Every Report Includes

Prompt Set
2026-07
Models
Claude 4, Gemini, GPT-5
Sample
4,218 queries
Period
30 days
Significance
p < 0.05

Every finding includes how it was produced, even if it can't be fully reproduced.

Confidence Distribution

90+■■■■■■42 findings
80+■■■18 findings
70+5 findings
60+2 findings
<601 findings

We don't hide low-confidence findings. We show the full distribution.

Critical Rule

If we ever publish "research" based on synthetic data without clear labeling, that credibility is nearly impossible to recover. We treat data integrity as our highest priority.

No Simulated Data in Public Reports

Seed data is used exclusively for development. No public report is ever published based on simulated or incomplete data.

Methodology Transparency

Every published finding includes methodology: number of queries, which AI models tested, time period covered, and significance criteria used.

Statistical Rigor

We require sufficient sample sizes before publishing. A single anomalous response is not a finding. Trends are confirmed across multiple queries and time periods.

Full Archive Access

Our raw data is verifiable. The AI Search Archive™ stores every response with full provenance: prompt, model, citations, entities, date.

Confidence Gates

Before any finding is published, it must pass evidence scoring, confidence evaluation, and editorial review. Low-confidence observations are flagged, never hidden.

Independent & Unbiased

We are an independent research center. We do not favor any AI model, platform, or company. Our only allegiance is to accurate, reproducible data.

Observatory DOI

OBS-2026-0042

Like academic DOI. Citable. Permanent. Never changes.

Permanent URLs

/research/2026/07/chatgpt-github-citations

Never /latest. Always permanent. This is an archive.

Three Data Modes

Development

Everything allowed, simulated data visible

Preview

Everything allowed, simulated data clearly labeled

Production

If isSimulated==true, API REFUSES. No option. Not even by mistake.

All public findings include full methodology, sample sizes, and confidence scores