{"schemaVersion":"maha-epistemic/1.0","evidencePolicyVersion":"mps/0.1","generatedAt":"2026-08-24T00:00:00.000Z","domain":{"slug":"mechanistic-interpretability","name":"Mechanistic interpretability","description":"Features, circuits, interventions, sparse representations, causal tests, and faithfulness criteria represented separately from explanatory confidence.","stressPoint":"A human-readable feature label, probe, attention pattern, or ablation effect does not by itself establish a complete or faithful causal explanation.","accent":"blue"},"lifecycle":{"status":"adversarial-pilot","foundationalTarget":30,"canonicalFactoryRecords":11,"outstandingFactoryRecords":19},"counts":{"graphRecords":30,"graphEdges":34,"publicCanonicalRecords":11,"withheldRecords":19},"records":[{"id":"urn:maha:record:mechanistic-interpretability-representation-probing-boundary","title":"Representation probing boundary","recordKind":"comparison","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/comparisons/mechanistic-interpretability-representation-probing-boundary","contentHash":"sha256:e43d7c00b47a6c78bb68667e5caf138a831208d102a9e8c8e013ef3a3746c3ff","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-representation-probing-boundary","scope":"Limited to Definitions, toy models, geometry, sparsity, and feature-interference experiments. in “Toy Models of Superposition”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Representation probing boundary does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"empirical-claim","sourceIds":["source-mechanistic-interpretability-superposition"],"statement":"The cited source supports treating representation probing boundary as a distinct comparison within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-superposition","url":"https://transformer-circuits.pub/2022/toy_model/index.html","title":"Toy Models of Superposition","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Nelson Elhage","Tristan Hume","Catherine Olsson","et al."],"boundary":"A toy-model mechanism does not establish that every feature in a production model has the same geometry or semantics.","publisher":"Transformer Circuits Thread","establishes":"The work develops toy models in which neural networks represent more features than available dimensions under specified sparsity conditions.","identifiers":[{"value":"https://transformer-circuits.pub/2022/toy_model/index.html","scheme":"url"}],"publishedAt":"2022-09-14","exactLocator":"Definitions, toy models, geometry, sparsity, and feature-interference experiments."}],"boundaries":["Representation probing boundary does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this representation probing boundary record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-induction-head-circuits","title":"Induction head circuits","recordKind":"concept","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/concepts/mechanistic-interpretability-induction-head-circuits","contentHash":"sha256:abdbf4922d70f2f43e85fe6cbaf7983d3355a151cafe9177c0660fdfccf66405","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-induction-head-circuits","scope":"Limited to Induction-head definition, previous-token heads, training dynamics, interventions, and model scope. in “In-context Learning and Induction Heads”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Induction head circuits does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-induction"],"statement":"The cited source supports treating induction head circuits as a distinct concept within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-induction","url":"https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html","title":"In-context Learning and Induction Heads","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Catherine Olsson","Nelson Elhage","Neel Nanda","et al."],"boundary":"Observed circuits in studied models do not establish a universal account of in-context learning.","publisher":"Transformer Circuits Thread","establishes":"The work reports circuits and interventions associated with induction-like behavior in specified transformer models.","identifiers":[{"value":"https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html","scheme":"url"}],"publishedAt":"2022-03-22","exactLocator":"Induction-head definition, previous-token heads, training dynamics, interventions, and model scope."}],"boundaries":["Induction head circuits does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this induction head circuits record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-neural-feature-superposition","title":"Neural feature superposition","recordKind":"concept","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/concepts/mechanistic-interpretability-neural-feature-superposition","contentHash":"sha256:e6cc01ac1e7583e26f9ccc1aa521459f050f2b92c08b21e1095f6fa4543754c5","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-neural-feature-superposition","scope":"Limited to Definitions, toy models, geometry, sparsity, and feature-interference experiments. in “Toy Models of Superposition”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Neural feature superposition does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-superposition"],"statement":"The cited source supports treating neural feature superposition as a distinct concept within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-superposition","url":"https://transformer-circuits.pub/2022/toy_model/index.html","title":"Toy Models of Superposition","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Nelson Elhage","Tristan Hume","Catherine Olsson","et al."],"boundary":"A toy-model mechanism does not establish that every feature in a production model has the same geometry or semantics.","publisher":"Transformer Circuits Thread","establishes":"The work develops toy models in which neural networks represent more features than available dimensions under specified sparsity conditions.","identifiers":[{"value":"https://transformer-circuits.pub/2022/toy_model/index.html","scheme":"url"}],"publishedAt":"2022-09-14","exactLocator":"Definitions, toy models, geometry, sparsity, and feature-interference experiments."}],"boundaries":["Neural feature superposition does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this neural feature superposition record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries","title":"Sparse autoencoder dictionaries","recordKind":"concept","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/concepts/mechanistic-interpretability-sparse-autoencoder-dictionaries","contentHash":"sha256:4685df47be76d62b550e5cab12a91f0fbaabc8c6db581c07bc660a59b8c081ca","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-sparse-autoencoder-dictionaries","scope":"Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-sae"],"statement":"The cited source supports treating sparse autoencoder dictionaries as a distinct concept within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-sae","url":"https://arxiv.org/abs/2309.08600","title":"Sparse Autoencoders Find Highly Interpretable Features in Language Models","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Hoagy Cunningham","Aidan Ewart","Logan Riggs","Robert Huben","Lee Sharkey"],"boundary":"Sparse features and human labels do not establish completeness, unique decomposition, or causal faithfulness.","publisher":"arXiv","establishes":"The paper trains sparse autoencoders on language-model activations and evaluates specified reconstruction, sparsity, and interpretability properties.","identifiers":[{"value":"https://arxiv.org/abs/2309.08600","scheme":"url"}],"publishedAt":"2023-09-15","exactLocator":"Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations."}],"boundaries":["Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this sparse autoencoder dictionaries record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-attention-pattern-evidence","title":"Attention pattern evidence","recordKind":"measurement","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/measurements/mechanistic-interpretability-attention-pattern-evidence","contentHash":"sha256:93e89d9f18dd46d8294b9a1877ae5fd3fc9b61f92143ce3e39f4e8d49f349a2c","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-attention-pattern-evidence","scope":"Limited to Induction-head definition, previous-token heads, training dynamics, interventions, and model scope. in “In-context Learning and Induction Heads”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Attention pattern evidence does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"observation","sourceIds":["source-mechanistic-interpretability-induction"],"statement":"The cited source supports treating attention pattern evidence as a distinct measurement within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-induction","url":"https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html","title":"In-context Learning and Induction Heads","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Catherine Olsson","Nelson Elhage","Neel Nanda","et al."],"boundary":"Observed circuits in studied models do not establish a universal account of in-context learning.","publisher":"Transformer Circuits Thread","establishes":"The work reports circuits and interventions associated with induction-like behavior in specified transformer models.","identifiers":[{"value":"https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html","scheme":"url"}],"publishedAt":"2022-03-22","exactLocator":"Induction-head definition, previous-token heads, training dynamics, interventions, and model scope."}],"boundaries":["Attention pattern evidence does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this attention pattern evidence record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-superposition-geometry","title":"Superposition geometry","recordKind":"measurement","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/measurements/mechanistic-interpretability-superposition-geometry","contentHash":"sha256:08c91d39d301182e33e89a3e76bcd9fd1a8da8ea42a0d34b0665963a2e0a3458","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-superposition-geometry","scope":"Limited to Definitions, toy models, geometry, sparsity, and feature-interference experiments. in “Toy Models of Superposition”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Superposition geometry does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"observation","sourceIds":["source-mechanistic-interpretability-superposition"],"statement":"The cited source supports treating superposition geometry as a distinct measurement within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-superposition","url":"https://transformer-circuits.pub/2022/toy_model/index.html","title":"Toy Models of Superposition","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Nelson Elhage","Tristan Hume","Catherine Olsson","et al."],"boundary":"A toy-model mechanism does not establish that every feature in a production model has the same geometry or semantics.","publisher":"Transformer Circuits Thread","establishes":"The work develops toy models in which neural networks represent more features than available dimensions under specified sparsity conditions.","identifiers":[{"value":"https://transformer-circuits.pub/2022/toy_model/index.html","scheme":"url"}],"publishedAt":"2022-09-14","exactLocator":"Definitions, toy models, geometry, sparsity, and feature-interference experiments."}],"boundaries":["Superposition geometry does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this superposition geometry record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-polysemantic-neurons","title":"Polysemantic neurons","recordKind":"mechanism","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/mechanisms/mechanistic-interpretability-polysemantic-neurons","contentHash":"sha256:a98f3991ac8186fdce269bfd35cd3d8823ea2705562f6b01f7cf8bf719654748","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-polysemantic-neurons","scope":"Limited to Definitions, toy models, geometry, sparsity, and feature-interference experiments. in “Toy Models of Superposition”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Polysemantic neurons does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-superposition"],"statement":"The cited source supports treating polysemantic neurons as a distinct mechanism within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-superposition","url":"https://transformer-circuits.pub/2022/toy_model/index.html","title":"Toy Models of Superposition","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Nelson Elhage","Tristan Hume","Catherine Olsson","et al."],"boundary":"A toy-model mechanism does not establish that every feature in a production model has the same geometry or semantics.","publisher":"Transformer Circuits Thread","establishes":"The work develops toy models in which neural networks represent more features than available dimensions under specified sparsity conditions.","identifiers":[{"value":"https://transformer-circuits.pub/2022/toy_model/index.html","scheme":"url"}],"publishedAt":"2022-09-14","exactLocator":"Definitions, toy models, geometry, sparsity, and feature-interference experiments."}],"boundaries":["Polysemantic neurons does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this polysemantic neurons record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-sae-encoder-decoder","title":"SAE encoder decoder","recordKind":"mechanism","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/mechanisms/mechanistic-interpretability-sae-encoder-decoder","contentHash":"sha256:79445a65081305ad694085c50fbdd859d0a04b3503ec55f5839a3e76d5576ee1","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-sae-encoder-decoder","scope":"Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"SAE encoder decoder does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-sae"],"statement":"The cited source supports treating sae encoder decoder as a distinct mechanism within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-sae","url":"https://arxiv.org/abs/2309.08600","title":"Sparse Autoencoders Find Highly Interpretable Features in Language Models","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Hoagy Cunningham","Aidan Ewart","Logan Riggs","Robert Huben","Lee Sharkey"],"boundary":"Sparse features and human labels do not establish completeness, unique decomposition, or causal faithfulness.","publisher":"arXiv","establishes":"The paper trains sparse autoencoders on language-model activations and evaluates specified reconstruction, sparsity, and interpretability properties.","identifiers":[{"value":"https://arxiv.org/abs/2309.08600","scheme":"url"}],"publishedAt":"2023-09-15","exactLocator":"Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations."}],"boundaries":["SAE encoder decoder does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this sae encoder decoder record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-causal-scrubbing","title":"Causal scrubbing","recordKind":"method","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/methods/mechanistic-interpretability-causal-scrubbing","contentHash":"sha256:4097a1b90c0c4e9908313d83025f6f229ef66fb4610c058164540d8363ef5cbe","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-causal-scrubbing","scope":"Limited to Method definition, correspondence, resampling interventions, examples, and limitations. in “Causal Scrubbing: a method for rigorously testing interpretability hypotheses”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Causal scrubbing does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-causal-scrubbing"],"statement":"The cited source supports treating causal scrubbing as a distinct method within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-causal-scrubbing","url":"https://www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing","title":"Causal Scrubbing: a method for rigorously testing interpretability hypotheses","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Lawrence Chan","Adrià Garriga-Alonso","Nicholas Goldowsky-Dill","et al."],"boundary":"Passing a declared test does not prove the hypothesis is unique, complete, or semantically correct.","publisher":"Alignment Research Center","establishes":"The work proposes resampling-based tests for whether a hypothesized computational graph preserves behavior under declared interventions.","identifiers":[{"value":"https://www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing","scheme":"url"}],"publishedAt":"2022-12-06","exactLocator":"Method definition, correspondence, resampling interventions, examples, and limitations."}],"boundaries":["Causal scrubbing does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this causal scrubbing record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-in-context-learning-circuits","title":"In context learning circuits","recordKind":"method","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/methods/mechanistic-interpretability-in-context-learning-circuits","contentHash":"sha256:5da6b19ec9d68ea868c2d4ff014d4f4c7a83dac77f81f37f9ee0172da2e05af1","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-in-context-learning-circuits","scope":"Limited to Induction-head definition, previous-token heads, training dynamics, interventions, and model scope. in “In-context Learning and Induction Heads”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"In context learning circuits does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-induction"],"statement":"The cited source supports treating in context learning circuits as a distinct method within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-induction","url":"https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html","title":"In-context Learning and Induction Heads","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Catherine Olsson","Nelson Elhage","Neel Nanda","et al."],"boundary":"Observed circuits in studied models do not establish a universal account of in-context learning.","publisher":"Transformer Circuits Thread","establishes":"The work reports circuits and interventions associated with induction-like behavior in specified transformer models.","identifiers":[{"value":"https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html","scheme":"url"}],"publishedAt":"2022-03-22","exactLocator":"Induction-head definition, previous-token heads, training dynamics, interventions, and model scope."}],"boundaries":["In context learning circuits does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this in context learning circuits record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]},{"id":"urn:maha:record:mechanistic-interpretability-toy-models-of-superposition","title":"Toy models of superposition","recordKind":"method","reviewState":"published-canonical","canonicalPath":"/knowledge/mechanistic-interpretability/methods/mechanistic-interpretability-toy-models-of-superposition","contentHash":"sha256:65e9b289d4e37ce598e80f3ea2a4abdb947ed9685cc0f77ee4013a6855fcbc75","claims":[{"id":"urn:maha:claim:mechanistic-interpretability-toy-models-of-superposition","scope":"Limited to Definitions, toy models, geometry, sparsity, and feature-interference experiments. in “Toy Models of Superposition”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Toy models of superposition does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-superposition"],"statement":"The cited source supports treating toy models of superposition as a distinct method within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-superposition","url":"https://transformer-circuits.pub/2022/toy_model/index.html","title":"Toy Models of Superposition","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Nelson Elhage","Tristan Hume","Catherine Olsson","et al."],"boundary":"A toy-model mechanism does not establish that every feature in a production model has the same geometry or semantics.","publisher":"Transformer Circuits Thread","establishes":"The work develops toy models in which neural networks represent more features than available dimensions under specified sparsity conditions.","identifiers":[{"value":"https://transformer-circuits.pub/2022/toy_model/index.html","scheme":"url"}],"publishedAt":"2022-09-14","exactLocator":"Definitions, toy models, geometry, sparsity, and feature-interference experiments."}],"boundaries":["Toy models of superposition does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record."],"prohibitedInferences":["Do not use this toy models of superposition record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.","Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract."]}],"withheldInventory":{"recordCount":19,"edgeCount":22,"recordKinds":{"comparison":5,"concept":3,"measurement":4,"mechanism":4,"method":3},"disclosure":"aggregate-only"},"boundary":"Only active canonical records are enumerated. Draft identifiers, titles, paths, claims, sources, and gate reasons remain private until canonical release."}