Evaluation and governance

Benchmarking, energy, and task equivalence

Compare systems only after aligning tasks, correctness, boundaries, amortization, and excluded costs.

hybridestablished research

Evidence status

Checked against 1 inspected source

One source was retrieved, identified and read, and the claims below are tied to specific passages at the scope those passages state. Each source also records what it cannot establish.

Rely on this page for

The specific claims that carry a cited passage, at the scope that passage states.

Working definition

A credible benchmark declares the task, dataset, preprocessing, accuracy or quality constraint, latency definition, measurement instrument, system boundary, idle power, host and interface costs, training or adaptation, repetitions, and uncertainty. Chip energy, wall-plug energy, biological metabolic cost, and laboratory support are different quantities and cannot share one unlabeled efficiency ranking.

Mechanism

  • Freeze task and correctness criteria.
  • Declare measurement and amortization boundaries.
  • Report paired performance, resource, and uncertainty metrics.

Measurements

  • Task quality
  • Latency and throughput
  • Energy, materials, labor, and support costs

Reproducibility controls

  • Version hardware, software, firmware, and analysis code.
  • Declare dataset, preprocessing, random seeds, and measurement boundary.
  • Report repeated runs, variation, exclusions, and failed trials.

Limits and failure modes

  • No benchmark covers general usefulness.
  • Cross-substrate totals require explicit accounting models.

Mathematical connection

Formal structure without substrate erasure

Predeclared task scoring

Score probabilistic or categorical outputs under a rule selected before benchmark results are known.

Inputs

  • Frozen predictions
  • Observed labels
  • Declared score

Outputs

  • Comparable task score
  • Uncertainty interval
  • Baseline difference

Limit: A proper score evaluates the declared forecasts; it does not make mismatched tasks, energy boundaries, or substrates equivalent.

Technical and governance sources

  1. [1]NeuroBench: Advancing Neuromorphic Computing Through Collaborative, Fair and Representative Benchmarking · National Institute of Standards and Technology

    Establishes: A community framework separating algorithm and system tracks and defining task, correctness, efficiency, and reporting procedures intended to make neuromorphic results more comparable and reproducible.

    Boundary: A benchmark ranks submitted systems on declared tasks and metrics. It does not prove general intelligence, biological equivalence, safety, usefulness outside the benchmark, or superiority under unreported host and data costs.

  2. [2]Taking Neuromorphic Computing to the Next Level with Loihi 2 · Intel Labs

    Establishes: An official description of the Loihi 2 research chip, its programmable neuron models, event-based communication, on-chip learning support, and the Lava software framework used to construct neuromorphic applications.

    Boundary: This is a vendor technical brief about a research platform. Performance and efficiency results remain workload-, configuration-, measurement-boundary-, and comparison-dependent and do not establish equivalence to biological intelligence.

  3. [3]Molecular computation of solutions to combinatorial problems · Science

    Establishes: A foundational experiment using molecular biology operations and DNA strands to encode and recover a solution to one small directed Hamiltonian-path instance.

    Boundary: The experiment demonstrates a bounded molecular computation. It does not establish practical general-purpose DNA computing, favorable end-to-end energy or latency, autonomous operation, or scalability beyond the reported instance.

Related concepts

Direct answer

  • A credible benchmark declares the task, dataset, preprocessing, accuracy or quality constraint, latency definition, measurement instrument, system boundary, idle power, host and interface costs, training or adaptation, repetitions, and uncertainty. Chip energy, wall-plug energy, biological metabolic cost, and laboratory support are different quantities and cannot share one unlabeled efficiency ranking.

Mechanism and method

  • Freeze task and correctness criteria.
  • Declare measurement and amortization boundaries.
  • Report paired performance, resource, and uncertainty metrics.

What is measured

  • Task quality
  • Latency and throughput
  • Energy, materials, labor, and support costs

Comparison: Artificial and spiking neural networks

  • A spike is not automatically a biological action potential.
  • Operation counts are architecture-specific.
  • Must not be read as: Do not infer biological realism, intelligence, or efficiency from the presence of spikes without matched task evidence and complete resource accounting.

Comparison: Silicon energy and biological metabolic cost

  • ATP or metabolic estimates are not wall-plug joules.
  • Culture support and readout may dominate.
  • Must not be read as: Do not claim biological energy superiority by comparing metabolism alone with a complete electronic system, or chip core energy with full laboratory support.

Comparison: Research demonstration and deployable system

  • A paper result is not an SLA.
  • Scaling samples or devices can change behavior.
  • Must not be read as: Do not market a simulation, chip prototype, cell culture, organoid, or closed-loop laboratory experiment as deployable computing without reliability, safety, scaling, and lifecycle evidence.

Limitations

  • No benchmark covers general usefulness.
  • Cross-substrate totals require explicit accounting models.

Boundaries declared by the cited sources

  • A benchmark ranks submitted systems on declared tasks and metrics. It does not prove general intelligence, biological equivalence, safety, usefulness outside the benchmark, or superiority under unreported host and data costs. (boundary declared by NeuroBench: Advancing Neuromorphic Computing Through Collaborative, Fair and Representative Benchmarking)
  • This is a vendor technical brief about a research platform. Performance and efficiency results remain workload-, configuration-, measurement-boundary-, and comparison-dependent and do not establish equivalence to biological intelligence. (boundary declared by Taking Neuromorphic Computing to the Next Level with Loihi 2)
  • The experiment demonstrates a bounded molecular computation. It does not establish practical general-purpose DNA computing, favorable end-to-end energy or latency, autonomous operation, or scalability beyond the reported instance. (boundary declared by Molecular computation of solutions to combinatorial problems)
  • A cross-platform benchmark study. It validates no vendor efficiency claim, declares no platform universally superior, and sets no absolute efficiency threshold, since its results depend on workload and configuration. It cannot support a general claim that neuromorphic hardware is more efficient than conventional hardware. (boundary declared by Benchmarking Neuromorphic Hardware and Its Energy Expenditure)

Bridge: Predeclared task scoring

  • Score probabilistic or categorical outputs under a rule selected before benchmark results are known.
  • Input: Frozen predictions
  • Input: Observed labels
  • Input: Declared score
  • Output: Comparable task score
  • Output: Uncertainty interval
  • Output: Baseline difference
  • Limit: A proper score evaluates the declared forecasts; it does not make mismatched tasks, energy boundaries, or substrates equivalent.

Related records