{"release":{"schemaVersion":"maha-epistemic-release/1.0","releaseId":"epirelease_210dc3c3474f4808b4c325b15074938d","releaseKind":"initial","status":"active","recordId":"urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries","domainSlug":"mechanistic-interpretability","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","canonicalPath":"/knowledge/mechanistic-interpretability/concepts/mechanistic-interpretability-sparse-autoencoder-dictionaries","canonicalVersion":"1.0.0","supersedesReleaseId":null,"approvals":[{"scope":"boundary-adequacy","reviewId":"epireview_02cc07c1a1d746579b5a2fd9700264b4","reviewSha256":"sha256:96e941605885c495ed2150dbe56cc751b178b454a3447bfe1b6f06ad99cc5177","reviewedAt":"2026-08-30T16:02:10.985Z","reviewerKind":"internal-editorial","reviewMethod":"Each criterion is recomputed from the exact record, its inspected alignment audit, source identity, exact locator, rights basis, claim scope, boundary, uncertainty, replication status, prohibited inferences, and revision digest."},{"scope":"domain-fidelity","reviewId":"epireview_5952700f64554eef93855b52f13f589b","reviewSha256":"sha256:2765bda172851d0f69833fd2ef6c4a8d22cb3d8baea0ac9e12928c0df8496286","reviewedAt":"2026-08-30T16:02:10.928Z","reviewerKind":"internal-editorial","reviewMethod":"Each criterion is recomputed from the exact record, its inspected alignment audit, source identity, exact locator, rights basis, claim scope, boundary, uncertainty, replication status, prohibited inferences, and revision digest."},{"scope":"rights-and-locator","reviewId":"epireview_5556acca7d574fefba3d4cd6807cef16","reviewSha256":"sha256:e2fdc98782562106d45a930696ab1091ab75db7f2e8f9ec86a65b8f2fb64fef7","reviewedAt":"2026-08-30T16:02:11.067Z","reviewerKind":"internal-editorial","reviewMethod":"Each criterion is recomputed from the exact record, its inspected alignment audit, source identity, exact locator, rights basis, claim scope, boundary, uncertainty, replication status, prohibited inferences, and revision digest."},{"scope":"source-fidelity","reviewId":"epireview_3ade585ee5244a7f9b901acc82541380","reviewSha256":"sha256:fc4101a9d5e72ad34ecf8ae50b3440d63f46e189c58a1ea28c82d3a68652ecf2","reviewedAt":"2026-08-30T16:02:10.856Z","reviewerKind":"internal-editorial","reviewMethod":"Each criterion is recomputed from the exact record, its inspected alignment audit, source identity, exact locator, rights basis, claim scope, boundary, uncertainty, replication status, prohibited inferences, and revision digest."}],"assuranceTier":"internally-reviewed-canonical","releaseAuthority":{"authoritySha256":"sha256:2a0161a3927ae8366af44bda6671b5249802843ff3895dc292c90e5466b88609","attribution":"withheld-by-consent"},"publicChangeSummary":"Initial canonical publication under the disclosed exact-revision internal editorial tier.","recordSha256":"sha256:4685df47be76d62b550e5cab12a91f0fbaabc8c6db581c07bc660a59b8c081ca","gateDecision":{"reasons":[],"recordId":"urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries","publicEligible":true,"evaluatedAgainst":"maha-epistemic/1.0"},"releasedAt":"2026-08-30T16:03:35.327Z","releaseSha256":"sha256:2fc4457ff5ecc5e62ab9c6880397c95355d8792449087eb195dfea09960e698e","withdrawal":null},"provenance":{"schemaVersion":"maha-epistemic/1.0","evidencePolicyVersion":"mps/0.1","recordId":"urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries","canonicalPath":"/knowledge/mechanistic-interpretability/concepts/mechanistic-interpretability-sparse-autoencoder-dictionaries","contentHash":"sha256:4685df47be76d62b550e5cab12a91f0fbaabc8c6db581c07bc660a59b8c081ca","generatedAt":"2026-08-30T16:03:35.327Z","publicationDecision":{"recordId":"urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries","publicEligible":true,"evaluatedAgainst":"maha-epistemic/1.0","reasons":[]},"claims":[{"id":"urn:maha:claim:mechanistic-interpretability-sparse-autoencoder-dictionaries","scope":"Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-sae"],"statement":"The cited source supports treating sparse autoencoder dictionaries as a distinct concept within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-sae","url":"https://arxiv.org/abs/2309.08600","title":"Sparse Autoencoders Find Highly Interpretable Features in Language Models","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Hoagy Cunningham","Aidan Ewart","Logan Riggs","Robert Huben","Lee Sharkey"],"boundary":"Sparse features and human labels do not establish completeness, unique decomposition, or causal faithfulness.","publisher":"arXiv","establishes":"The paper trains sparse autoencoders on language-model activations and evaluates specified reconstruction, sparsity, and interpretability properties.","identifiers":[{"value":"https://arxiv.org/abs/2309.08600","scheme":"url"}],"publishedAt":"2023-09-15","exactLocator":"Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations."}],"reviewEvents":[{"reviewId":"epireview_3ade585ee5244a7f9b901acc82541380","scope":"source-fidelity","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","reviewedAt":"2026-08-30T16:02:10.856Z","verdict":"approve","supersedesReviewId":null},{"reviewId":"epireview_5952700f64554eef93855b52f13f589b","scope":"domain-fidelity","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","reviewedAt":"2026-08-30T16:02:10.928Z","verdict":"approve","supersedesReviewId":null},{"reviewId":"epireview_02cc07c1a1d746579b5a2fd9700264b4","scope":"boundary-adequacy","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","reviewedAt":"2026-08-30T16:02:10.985Z","verdict":"approve","supersedesReviewId":null},{"reviewId":"epireview_5556acca7d574fefba3d4cd6807cef16","scope":"rights-and-locator","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","reviewedAt":"2026-08-30T16:02:11.067Z","verdict":"approve","supersedesReviewId":null}]},"privacyBoundary":"Operational actor fingerprints, bearer credentials, private reviewer profiles, affiliations, conflicts, and non-consented authority identity fields are excluded."}