Site search

Find architecture, research, and terms

Start typing to search the editorial index.

Editorial guidanceAdvanced

Reproducibility, Evidence, and Open Questions

Replay levels, research-grade evidence packages, claim registers, negative results, and an open agenda for stochastic tool-using runtimes.

Replay levels

Exact replay

Inputs, versions, environment, state, and outputs match within a declared deterministic boundary.

Functionally equivalent replay

The same objective and policy outcome are reached although token-level or timing details differ.

Partial replay

Only selected model, tool, or policy steps can be reconstructed.

Unavailable replay

Required external state, model version, source snapshot, or side-effect history is absent.

Research-grade evidence package

  • Request and interpreted objective.
  • Runtime profile, model route/version, and prompt-contract version.
  • Context source IDs/hashes, memory scopes, tool/schema versions, and policy decisions.
  • Approvals, checkpoints, artifacts/hashes, errors, recovery, and unresolved uncertainty.
  • Random seed where available, environment/dependency versions, and UTC timestamps.

Evidence should minimize sensitive content while preserving the facts required for review. OpenTelemetry conventions can support correlation but do not replace product-level evidence semantics. Source: OpenTelemetry GenAI conventions

Claim register

Field Purpose
Claim Specific, falsifiable statement rather than a broad narrative.
Source type Foundational work, formal proposal, prototype, benchmark, production evidence, or editorial inference.
Evidence level Current position on the evidence ladder.
Counterevidence Known results or interpretations that weaken the claim.
Limitations Scope, assumptions, missing data, and boundary conditions.
Status and reviewed date Active, downgraded, deferred, rejected, or superseded with UTC review date.

Negative and null results

Publish adaptation that did not improve outcomes, structural edits that increased cost, critic loops that failed, no-op thresholds that were too conservative, and benchmark tasks with inconclusive results. Negative evidence is necessary to distinguish a reusable mechanism from selective reporting.

Open research agenda

  • A formal, falsifiable definition of teleodynamic runtime behavior.
  • Safe goal revision and resource variables that generalize across workloads.
  • Independent evaluators and causal attribution for recovery.
  • Operator control under self-maintenance and long-horizon drift.
  • Benchmark portability and evidence interoperability.
  • Criteria that separate useful metaphor from implemented mechanism.

Source record

References

Suggest a correction
  1. National Institute of Standards and Technology. NIST. Published 2023-01-26; last reviewed 2026-06-20 UTC. Government framework.

  2. National Institute of Standards and Technology. NIST. Published 2024-07-26; last reviewed 2026-06-20 UTC. Government profile.

  3. OpenTelemetry project. Cloud Native Computing Foundation. Published Current specification repository; last reviewed 2026-06-24 UTC. Official specification.

  4. Enrique ter Horst and Juan Diego Zambrano. arXiv. Published 2026-03-11; last reviewed 2026-06-24 UTC. Research preprint.