DTDeepTrailJudge mode

OpenAI WebMCP Challenge · 60-second start

Watch a browser agent challenge a decision inside the same evidence workspace a human can inspect.

DeepTrail turns transient web research into shared state: questions, sources, claims, evidence links, counterarguments, confidence changes, research debt, and a draft decision. WebMCP is the collaboration layer—not a wrapper around a chatbot.

Open workspace

Production readiness

Browser + WebMCP release checks

0/7 ready

Secure context

Checking

Checking whether this origin is secure.

WebMCP browser capability

Checking

Checking document.modelContext.

Origin isolation header

Checking

Checking Origin-Agent-Cluster.

Tools permissions policy

Checking

Checking Permissions-Policy.

Content type protection

Checking

Checking X-Content-Type-Options.

Referrer protection

Checking

Checking Referrer-Policy.

Local research storage

Checking

Checking IndexedDB.

What to do

Three steps for the strongest demo

  1. 1

    Load the seeded investigation

    It starts with sourced claims, a visible counterargument, open production-verification gaps, a comparison, and a deliberately draft decision.

  2. 2

    Give the browser agent the prompt below

    The agent reads stable IDs and structured state through WebMCP, searches the live web, and writes new evidence back through single-purpose tools.

  3. 3

    Watch the belief state move

    Look for the evidence graph, new counterevidence, confidence history, Research Debt, actor-attributed activity, and the decision remaining draft unless the evidence earns finality.

Exact judge prompt

Adversarial research, not agreeable summarization

Read the active DeepTrail investigation through WebMCP. Treat the current draft decision as a hypothesis, not a conclusion. First inspect the open production-verification questions and existing evidence. Then search for the strongest credible evidence that could falsify or materially qualify the current hosting recommendation. Add any new source with provenance, add or refine the relevant claim, link the evidence, record a counterargument if warranted, and update confidence only if the evidence justifies a change. Refresh research gaps at the end. Do not manufacture disagreement.

Why it fits the challenge

Built around the four judging criteria

WebMCP leverage

Shared application state

State-aware WebMCP tools let the agent inspect and mutate the same visible research objects the human edits.

Execution

Complete local-first product

IndexedDB persistence, provenance, recovery, strict validation, accessibility, and automated regression coverage make the demo resilient.

Potential impact

Better decisions, not more tabs

DeepTrail targets repeated research work where people need to understand why an answer is credible and what remains uncertain.

Creativity + ambition

Attack the conclusion

Falsification criteria, an evidence graph, confidence history, and deterministic Research Debt turn critical thinking into inspectable UI.

Judge demo