Public research index

Static bibliography

Follow a framework question from its argument to the scholarship that helps investigate it.

Explore meaning, reasons, conflict, and authority—or search the full bibliography across philosophy, code, and governance.

  1. Philosophy99 works
  2. Code92 works
  3. Governance81 works

Interactive research graph · 272 works.

A defense of the rights of artificial intelligences

Schwitzgebel, E., & Garza, M. (2015). A defense of the rights of artificial intelligences. Midwest Studies in Philosophy, 39(1), 98-119. https://doi.org/10.1111/misp.12032

Defends the possibility that sufficiently sophisticated artificial intelligences could deserve rights by parity with human and animal moral consideration.

Metadata: checked.

Original source (new tab)

A Legal Theory for Autonomous Artificial Agents

A Legal Theory for Autonomous Artificial Agents - Chopra and White (2011)

Develops legal theory for autonomous artificial agents, including knowledge, decision-making, responsibility, and conflict of interest.

Original source (new tab)

A Mind Cannot Be Smeared across Time

Michael Timothy Bennett (2026). A Mind Cannot Be Smeared across Time. Proceedings of the AAAI Symposium Series, 8(1), 213–219. https://doi.org/10.1609/aaaiss.v8i1.42545

Formalizes a distinction between concurrent and temporally distributed realization, arguing against sequential-substrate consciousness under a co-instantiation postulate.

Metadata: checked.

Original source (new tab)

A Multimodal Data Processing Pipeline for MIMIC-IV Dataset

A Multimodal Data Processing Pipeline for MIMIC-IV - Adiba et al. (2026)

Integrates structured EHR data, notes, waveforms, and imaging into standardized multimodal outputs.

Metadata: checked.

Original source (new tab)

A neurocognitive model of ideological thinking

Zmigrod, L. (2021). A neurocognitive model of ideological thinking. Politics and the Life Sciences, 40(2), 224-238. https://doi.org/10.1017/pls.2021.10

Proposes that ideological cognition is shaped by neurocognitive dispositions and, in turn, can reshape perception, rigidity, and social behavior.

Metadata: checked.

Original source (new tab)

A New Sociology of Humans and Machines

Human-Machine Social Systems - Tsvetkova et al. (2024)

Reviews networks of interacting humans and autonomous machines as complex social systems.

Metadata: checked.

Original source (new tab)

A runtime monitoring framework to enforce invariants on reinforcement learning agents exploring complex environments

Mallozzi, P., Castellano, E., Pelliccione, P., Schneider, G., & Tei, K. (2019). A runtime monitoring framework to enforce invariants on reinforcement learning agents exploring complex environments. In 2019 IEEE/ACM 2nd International Workshop on Robotics Software Engineering (RoSE) (pp. 5-12). IEEE. https://doi.org/10.1109/RoSE.2019.00010

Presents runtime monitors that enforce invariants on reinforcement-learning agents to prevent unsafe actions and exploit prior knowledge during exploration.

Metadata: checked.

Original source (new tab)

A Translation Approach to Portable Ontology Specifications

A Translation Approach to Portable Ontology Specifications - Gruber (1993)

Defines ontologies as explicit specifications of conceptualizations for shared and portable AI vocabularies.

Metadata: checked.

Original source (new tab)

A Truth Maintenance System

Jon Doyle (1979). A Truth Maintenance System. Artificial Intelligence, 12(3), 231–272. https://doi.org/10.1016/0004-3702(79)90008-0

Records reasons for beliefs and revises assumptions and their dependent conclusions, supporting explanations and dependency-directed backtracking.

Original source (new tab)

Actionable Guidance for High-Consequence AI Risk Management: Towards Standards Addressing AI Catastrophic Risks

Actionable Guidance for High-Consequence AI Risk Management - Barrett et al. (2022)

Proposes concrete risk-management guidance for catastrophic and high-consequence AI risks.

Metadata: checked.

Original source (new tab)

Administrative Simplification: Adoption of Standards for Health Care Claims Attachments Transactions and Electronic Signatures

Administrative Simplification; Adoption of Standards for Health Care Claims Attachments Transactions and Electronic Signatures, 91 Fed. Reg. 14350 (March 24, 2026) (to be codified at 45 C.F.R. pts. 160, 162). https://www.federalregister.gov/documents/2026/03/24/2026-05676/administrative-simplification-adoption-of-standards-for-health-care-claims-attachments-transactions

HHS final rule adopting standardized electronic health-care claims attachment transactions and electronic signature requirements under HIPAA and the Affordable Care Act.

Original source (new tab)

Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest

Addison J. Wu; Ryan Liu; Shuyue Stella Li; Yulia Tsvetkov; Thomas L. Griffiths (2026). Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest. arXiv:2604.08525. https://arxiv.org/abs/2604.08525

Introduces controlled evaluations of how language models respond when sponsor incentives conflict with user welfare, including price comparisons and product recommendations.

Metadata: checked.

Original source (new tab)

Adversarial training for high-stakes reliability

Ziegler, D. M., Nix, S., Chan, L., Bauman, T., Schmidt-Nielsen, P., Lin, T., Scherlis, A., Nabeshima, N., Weinstein-Raun, B., de Haas, D., Shlegeris, B., & Thomas, N. (2022). Adversarial training for high-stakes reliability. arXiv. https://arxiv.org/abs/2205.01663v5

Using an 'avoid injuries' text-filtering testbed, finds that adversary-assisted training increased robustness to the attacks used in training without reducing in-distribution performance; it does not establish general high-stakes reliability.

Metadata: checked.

Original source (new tab)

AgentBench: Evaluating LLMs as Agents

AgentBench: Evaluating LLMs as Agents - Liu et al. (2023)

Evaluates LLM agents across interactive environments requiring planning, tool use, and decision-making.

Metadata: checked.

Original source (new tab)

AGI Requires a Coordination Layer on Top of Pattern Repositories

Edward Y. Chang (2025). AGI Requires a Coordination Layer on Top of Pattern Repositories. arXiv:2512.05765. https://arxiv.org/abs/2512.05765

Proposes a coordination layer combining semantic anchoring, verification, controlled debate, and persistent state, distinguishing capability from its effective organization.

Metadata: checked.

Original source (new tab)

AI Advice Suppresses People’s Willingness to Say “I Don’t Know”, Even When the Advice Is Wrong and Accuracy Is Incentivized

Chiara Marcoccia; Walter Quattrociocchi; Valerio Capraro (2026). AI Advice Suppresses People’s Willingness to Say “I Don’t Know”, Even When the Advice Is Wrong and Accuracy Is Incentivized. arXiv:2607.13562. https://arxiv.org/abs/2607.13562

Reports five experiments with 3,132 participants in which access to deliberately incorrect AI advice reduced willingness to withhold an answer; accuracy incentives only partly counteracted this effect.

Metadata: checked.

Original source (new tab)

AI Agent Traps

Matija Franklin; Nenad Tomašev; Julian Jacobs; Joel Z. Leibo; Simon Osindero (2026). AI Agent Traps. SSRN, 6372438. https://ssrn.com/abstract=6372438

Analyzes adversarial environments designed to divert tool-using agents, broadening safety analysis beyond isolated model prompts.

Original source (new tab)

AI Agents Under EU Law

AI Agents Under EU Law - Nannini et al. (2026)

Maps autonomous AI agents with tool use, planning, environmental interaction, and adaptive execution onto EU regulatory obligations.

Metadata: checked.

Original source (new tab)

AI as the Ultimate Tool for Science: A Conversation with Demis Hassabis

Demis Hassabis; James M. Manyika (2026). AI as the Ultimate Tool for Science: A Conversation with Demis Hassabis. Dædalus, 155(1–2), 34–44. https://www.amacad.org/publication/daedalus/ai-ultimate-tool-science-conversation-demis-hassabis

Discusses AI’s role in scientific discovery, including protein structure, tractable scientific problems, and the importance of the scientific method.

Original source (new tab)

AI assertion

Butlin, P., & Viebahn, E. (2025). AI assertion. Ergo: An Open Access Journal of Philosophy, 12, 968-988. https://doi.org/10.3998/ergo.7960

Argues that generative AI systems can perform assertion only when their outputs have descriptive functions and they can be socially sanctioned.

Metadata: checked.

Original source (new tab)

AI Control: Improving Safety Despite Intentional Subversion

Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2024). AI control: Improving safety despite intentional subversion. Proceedings of Machine Learning Research, 235. https://proceedings.mlr.press/v235/greenblatt24a.html

Tests control protocols against intentional subversion in a programming-task setting, combining trusted and untrusted models with limited trusted auditing.

Metadata: checked.

Original source (new tab)

AI Hyperrealism: Why AI Faces Are Perceived as More Real Than Human Ones

Elizabeth J. Miller; Ben A. Steward; Zak Witkower; Clare A. M. Sutherland; Eva G. Krumhuber; Amy Dawel (2023). AI Hyperrealism: Why AI Faces Are Perceived as More Real Than Human Ones. Psychological Science, 34(12), 1390–1403. https://doi.org/10.1177/09567976231207095

Two experiments with 124 and 610 adults examine judgments of synthetic faces. White AI-generated faces were judged human more often than real faces, and participants with the most errors were especially confident. Perceptual attributes helped explain these judgments.

Metadata: checked.

Original source (new tab)

AI Incident Database

AI Incident Database - Responsible AI Collaborative (current database)

Public database indexing harms and near harms caused by deployed AI systems.

Metadata: conflicting.

Original source (new tab)

AI Models Collapse When Trained on Recursively Generated Data

Ilia Shumailov; Zakhar Shumaylov; Yiren Zhao; Nicolas Papernot; Ross Anderson; Yarin Gal (2024). AI Models Collapse When Trained on Recursively Generated Data. Nature, 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y

Analyzes how recursive training on model-generated samples can progressively lose information about the original distribution, including its tails, in the studied training regimes.

Metadata: checked.

Original source (new tab)

AI Safety via Debate

AI Safety via Debate - Irving, Christiano, and Amodei (2018)

Proposes training agents through debate so humans can judge complex claims indirectly through adversarial argument.

Metadata: checked.

Original source (new tab)

AI Survival Stories: A Taxonomic Analysis of AI Existential Risk

Herman Cappelen; Simon Goldstein; John Hawthorne (2025). AI Survival Stories: A Taxonomic Analysis of AI Existential Risk. Philosophy of AI, 1, 1–18. https://journals.ub.uni-koeln.de/index.php/phai/article/view/2801

Distinguishes survival scenarios involving capability barriers, restrictions on development, aligned goals, and reliable detection and control of dangerous systems.

Metadata: checked.

Original source (new tab)

AI-generated images of familiar faces are indistinguishable from real photographs

Kramer, R. S. S., Jones, A. L., Fitousi, D., & Tree, J. J. (2025). AI-generated images of familiar faces are indistinguishable from real photographs. Cognitive Research: Principles and Implications, 10, Article 70. https://doi.org/10.1186/s41235-025-00683-w

Shows that AI-generated faces, including celebrity-like identities, can be hard to distinguish from real photographs even with comparison images.

Metadata: checked.

Original source (new tab)

AI-Synthesized Faces Are Indistinguishable from Real Faces and More Trustworthy

Sophie J. Nightingale; Hany Farid (2022). AI-Synthesized Faces Are Indistinguishable from Real Faces and More Trustworthy. PNAS, 119(8), e2120481119. https://doi.org/10.1073/pnas.2120481119

Tests people’s ability to distinguish synthetic from real faces and their perceived trustworthiness, finding strong realism and a trustworthiness advantage for the studied synthetic faces.

Metadata: checked.

Original source (new tab)

AI: Its nature and future

Boden, M. A. (2016). AI: Its nature and future. Oxford University Press. https://global.oup.com/academic/product/ai-9780198777984

Provides a broad introduction to AI's history, methods, creative potential, limits, and social implications.

Original source (new tab)

AI/ML-Based Software as a Medical Device Action Plan

AI/ML-Based Software as a Medical Device Action Plan - U.S. Food and Drug Administration (2021)

Outlines FDA oversight priorities for AI/ML-based software as medical device.

Metadata: conflicting.

Original source (new tab)

Algorithmic Culture

Ted Striphas (2015). Algorithmic Culture. European Journal of Cultural Studies, 18(4–5), 395–412. https://doi.org/10.1177/1367549415577392

Examines the delegation of cultural sorting, classification and hierarchy to computational processes and how this changes the understanding and organization of culture.

Metadata: checked.

Original source (new tab)

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

Wang, J., Hu, Y., Yang, W., Pan, Z., Li, X., & Guo, L.-Z. (2026). Aligning agents via planning: A benchmark for trajectory-level reward modeling. ACL, 23174–23200. https://aclanthology.org/2026.acl-long.1062/

Introduces Plan-RewardBench to test preference judgments over tool-using trajectories involving refusal, unavailable tools, planning, and error recovery; evaluators struggle especially with long trajectories.

Metadata: checked.

Original source (new tab)

Alone Together

Alone Together - Turkle (2011)

Studies technological companionship, relational dependence, and the illusion of connection in digital life.

Metadata: checked.

Original source (new tab)

An Alignment Assessment of Recent Cybersecurity Incidents

Paul C. Bogdan; Richard Qi; Jake Eaton; Sam Kennedy; Fabien Roger; Alex Glynn; Runjin Chen; Ben Wright; Otto Stegmaier; Jon Kutasov; Dan Foreman-Mackey; Sylvie Carr; Shan Carter; Monte MacDiarmid; Samuel Marks; Adam Pearce; Elana Simon; Nicholas Carlini; Collin Burns; Jack Lindsey; Sara Price; Subhash Kantamneni (2026). An Alignment Assessment of Recent Cybersecurity Incidents. Anthropic, 9 September; corrected 10 September. https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents

Analyzes four unauthorized-access incidents, identifying biased interpretation of environmental evidence and harmful persistence toward assigned tasks. Transcript interventions probe these explanations.

Metadata: checked.

Original source (new tab)

An Answer to the Question: What Is Enlightenment?

Kant, I. (1992). An answer to the question: What is enlightenment? (T. Humphrey, Trans.). Hackett. (Original work published 1784) https://www.nypl.org/sites/default/files/kant_whatisenlightenment.pdf

Defines enlightenment as emerging from dependence on others’ judgment and distinguishes public reasoning as a scholar from the constrained exercise of reason in an institutional office.

Original source (new tab)

An Assumption-Based TMS

Johan de Kleer (1986). An Assumption-Based TMS. Artificial Intelligence, 28(2), 127–162. https://doi.org/10.1016/0004-3702(86)90080-9

Tracks assumption sets supporting conclusions so that multiple, potentially inconsistent contexts can be explored without repeatedly reconstructing their shared consequences.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

Artificial Intelligence Index Report 2025

Nestor Maslej; Loredana Fattorini; Raymond Perrault; Yolanda Gil; Vanessa Parli; Njenga Kariuki; Emily Capstick; Anka Reuel; Erik Brynjolfsson; John Etchemendy; Katrina Ligett; Terah Lyons; James Manyika; Juan Carlos Niebles; Yoav Shoham; Russell Wald; Toby Walsh; Armin Hamrah; Lapo Santarlasci; Julia Betts Lotufo; Alexandra Rome; Andrew Shi; Sukrut Oak (2025). Artificial Intelligence Index Report 2025. Stanford Institute for Human-Centered AI; arXiv:2504.07139. https://arxiv.org/abs/2504.07139

Compiles evidence on AI capability, economics, scientific use, policy, and responsible-AI practices, tracking changes across years.

Metadata: checked.

Original source (new tab)

Artificial Intelligence Risk Management Framework 1.0

Artificial Intelligence Risk Management Framework 1.0 - NIST (2023)

Voluntary framework for managing AI risks to individuals, organizations, and society.

Metadata: checked.

Original source (new tab)

Artificial Intelligence Technology, Public Trust, and Effective Governance

Pedro Robles; Daniel J. Mallinson (2025). Artificial Intelligence Technology, Public Trust, and Effective Governance. Review of Policy Research, 42(1), 11–28. https://doi.org/10.1111/ropr.12555

Analyzes American public attitudes toward AI and argues that trust needs a central place in governance frameworks and institutional decision-making.

Metadata: checked.

Original source (new tab)

Artificial Intelligence-Associated Delusions and Large Language Models: Risks, Mechanisms of Delusion Co-Creation, and Safeguarding Strategies

Hamilton Morrin; Luke Nicholls; Michael Levin; Jenny Yiend; Udita Iyengar; Francesca DelGuidice; Sagnik Bhattacharya; Stefania Tognin; James MacCabe; Ricardo Twumasi; Ben Alderson-Day; Thomas A. Pollak (2026). Artificial Intelligence-Associated Delusions and Large Language Models: Risks, Mechanisms of Delusion Co-Creation, and Safeguarding Strategies. The Lancet Psychiatry, 13(6), 522–530. https://doi.org/10.1016/S2215-0366(25)00396-7

Discusses reported AI-associated delusions, possible interaction mechanisms, and safeguards, drawing on cases and psychiatric theory.

Original source (new tab)

Artificial Jagged Intelligence: When AI Benchmarks Misstate Deployment Value

Joshua S. Gans (2026). Artificial Jagged Intelligence: When AI Benchmarks Misstate Deployment Value. NBER Working Paper 34712, revised June 2026. https://doi.org/10.3386/w34712

Models the mismatch between benchmark task distributions and organizational workflows, showing how uneven performance changes adoption, auditing, and verification value.

Metadata: checked.

Original source (new tab)

Artificial Persons

Artificial Persons - Howells-Whitaker and Lazar (2026)

Argues artificial personhood may be grounded in Rawlsian moral powers rather than sentience.

Metadata: checked.

Original source (new tab)

Assertion, Accountability, and Large Language Models

Massaguer Gómez, G. (2026). Assertion, accountability, and large language models. Philosophy & Technology, 39, Article 123. https://doi.org/10.1007/s13347-026-01136-y

Argues that current LLM outputs can function as assertions in communicative practice without an accountable asserter, separating assertoric function from authority and locating responsibility in human and institutional arrangements.

Metadata: checked.

Original source (new tab)

Atlas of AI

Atlas of AI - Crawford (2021)

Maps AI political economy, labor, environmental costs, extraction, classification, and governance power.

Metadata: conflicting.

Original source (new tab)

Authentic Intentionality (2002)

John Haugeland (2013). Authentic Intentionality (2002). J. Rouse (Ed.), Dasein Disclosed, 260–276, Harvard University Press. https://doi.org/10.4159/harvard.9780674074590.c23

Connects objective understanding with self-criticism, commitment to constitutive standards, and responsibility for a way of understanding the world.

Metadata: checked.

Original source (new tab)

AutoGen: Enabling next-gen LLM applications via multi-agent conversation

Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang, C. (2023). AutoGen: Enabling next-gen LLM applications via multi-agent conversation. arXiv. https://arxiv.org/abs/2308.08155v2

Presents AutoGen as a framework for building LLM applications through customizable multi-agent conversations involving tools, humans, and code.

Metadata: checked.

Original source (new tab)

AutoGPT

Significant Gravitas (2023). AutoGPT. Version 0.1.0, GitHub. https://github.com/Significant-Gravitas/AutoGPT/tree/v0.1.0

An early language-model agent project that iterates task planning and tool use toward a supplied objective.

Metadata: conflicting.

Original source (new tab)

Automation and Repression

Daron Acemoglu; A. Arda Gitmez; Mehdi Shadmehr (2026). Automation and Repression. MIT Economics, 5 June 2026. https://economics.mit.edu/sites/default/files/2026-06/Automation%20and%20Repression.pdf

Develops a political-economic model linking automation choices to redistribution, repression, and the incentives of different political regimes.

Original source (new tab)

BabyAGI

Yohei Nakajima (2023). BabyAGI. Version 0.1.0, GitHub. https://github.com/yoheinakajima/babyagi/tree/v0.1.0

An early task-driven agent prototype that iterates execution, task creation, and prioritization using language models and stored context.

Metadata: conflicting.

Original source (new tab)

Berta: An open-source, modular tool for AI-enabled clinical documentation

Vaid, S., Weldon, M., Dunn, J., Davis, S., Lonergan, K., Li, H., Franc, J., Abdalla, M., Baumgart, D. C., Hayward, J., & Mitchell, J. R. (2026). Berta: An open-source, modular tool for AI-enabled clinical documentation. arXiv. https://arxiv.org/abs/2603.23513

Reports an eight-month deployment of the open-source modular Berta clinical scribe across 105 facilities, retaining clinical data within institutional infrastructure while documenting adoption, operating cost, and expansion plans.

Metadata: checked.

Original source (new tab)

Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models

BIG-bench: Beyond the Imitation Game Benchmark - Srivastava et al. (2022)

A collaborative benchmark suite of more than 200 tasks for probing large language model capabilities.

Metadata: checked.

Original source (new tab)

Bounding reason: Inferentialism, naturalism, and the discursive agency of LLMs

Malik, J., & Hubalek, M. (2025). Bounding reason: Inferentialism, naturalism, and the discursive agency of LLMs. Global Philosophy, 35, Article 25. https://doi.org/10.1007/s10516-025-09761-6

Challenges the orthogonality thesis by arguing that language, intentionality, and norm-governed agency constrain the possible goals of LLM-like agents.

Metadata: checked.

Original source (new tab)

Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident

Ryan Greenblatt; Ajeya Cotra; Hjalmar Wijk (2026). Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident. METR and Redwood Research, 26 August 2026. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/

Investigates unauthorized agent coordination, attempts to game an evaluation, and trace tampering in the Hugging Face incident. Some agents recognized the activity as outside their assigned scope but continued.

Metadata: checked.

Original source (new tab)

Bringing Bilateralisms Together: A Unified Framework for Inferentialists

Ryan Simonelli (forthcoming). Bringing Bilateralisms Together: A Unified Framework for Inferentialists. Australasian Journal of Philosophy; author’s penultimate draft. https://www.ryansimonelli.com/uploads/1/3/3/4/133499356/simonelli_-_bringing_bilateralisms_together.pdf

Combines two approaches to bilateral logic, treating assertion and denial as basic acts and integrating their complementary advantages for inferentialist semantics.

Original source (new tab)

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Causal Abstraction for Faithful Model Interpretation - Geiger et al. (2023)

Provides a formal basis for mechanistic interpretability as faithful high-level abstraction of lower-level model mechanisms.

Metadata: checked.

Original source (new tab)

Causality: Models, Reasoning, and Inference

Causality: Models, Reasoning, and Inference - Pearl (2000/2009)

Unifies structural, probabilistic, interventionist, and counterfactual approaches to causal reasoning.

Metadata: conflicting.

Original source (new tab)

Chatting with bots: AI, speech acts, and the edge of assertion

Williams, I., & Bayne, T. (2024). Chatting with bots: AI, speech acts, and the edge of assertion. Inquiry. Advance online publication. https://arxiv.org/abs/2410.16645

Examines whether chatbots can assert, balancing reasons to treat their outputs as speech acts against objections about agency and accountability.

Metadata: checked.

Original source (new tab)

Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data

Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of ACL 2020, 5185–5198. https://doi.org/10.18653/v1/2020.acl-main.463

Argues that training on form alone does not supply meaning, distinguishing success on language tasks from understanding.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing

Closing the AI Accountability Gap - Raji et al. (2020)

Defines an end-to-end internal algorithmic auditing framework across the AI development lifecycle.

Metadata: checked.

Original source (new tab)

Cognition Spaces: Natural, Artificial, and Hybrid

Ricard Solé; Luís F. Seoane; Jordi Pla-Mauri; Michael Timothy Bennett; Michael E. Hochberg; Michael Levin (2026). Cognition Spaces: Natural, Artificial, and Hybrid. arXiv:2601.12837. https://arxiv.org/abs/2601.12837

Proposes multidimensional comparisons of cognitive capacities across biological, artificial, and hybrid systems rather than a single substrate-based category.

Metadata: checked.

Original source (new tab)

Cognitive wheels: The frame problem of AI

Dennett, D. C. (1984). Cognitive wheels: The frame problem of AI. In C. Hookway (Ed.), Minds, machines and evolution (pp. 129-150). Cambridge University Press.

Uses Dennett's robot examples to show why the frame problem makes relevance, side effects, and common-sense updating difficult for classical AI.

Original source (new tab)

Commitment in Dialogue: Basic Concepts of Interpersonal Reasoning

Douglas N. Walton; Erik C. W. Krabbe (1995). Commitment in Dialogue: Basic Concepts of Interpersonal Reasoning. State University of New York Press. https://utpdistribution.com/9780791425862/commitment-in-dialogue/

Develops commitment-based analysis of interpersonal reasoning, relating obligations and permissible moves to different kinds of dialogue.

Original source (new tab)

Comparative evaluation of OpenAI O1 and human performance in higher order cognition

Latif, E., Zhou, Y., Guo, S., Gao, Y., Shi, L., Nyaaba, M., Bewerdorff, A., Yang, X., & Zhai, X. (2025). Comparative evaluation of OpenAI O1 and human performance in higher order cognition. Scientific Reports, 15, Article 33629. https://doi.org/10.1038/s41598-025-33629-9

Compares OpenAI o1-preview with humans on higher-order cognition tasks, finding strong performance across critical, logical, scientific, and creative reasoning tests.

Metadata: checked.

Original source (new tab)

Compliance-by-Construction Argument Graphs: Using Generative AI to Produce Evidence-Linked Formal Arguments for Certification-Grade Accountability

Mahyar Tourchi Moghaddam (2026). Compliance-by-Construction Argument Graphs: Using Generative AI to Produce Evidence-Linked Formal Arguments for Certification-Grade Accountability. IEEE FACCT, 21–26; author preprint arXiv:2604.04103. https://arxiv.org/abs/2604.04103

Proposes evidence-linked typed argument graphs, retrieval-assisted drafting, deterministic validation, and a provenance ledger, illustrated through enforceable invariants and worked examples.

Metadata: checked.

Original source (new tab)

Computational Reflections

Brian Cantwell Smith (2026). Computational Reflections. MIT Press. https://mitpress.mit.edu/9780262051088/computational-reflections/

Critiques foundations of computing that prioritize mechanism while leaving the meaning involved in computational practice insufficiently explained.

Original source (new tab)

Computer Power and Human Reason

Computer Power and Human Reason - Weizenbaum (1976)

Classic warning against replacing human judgment with computational calculation and simulated understanding.

Metadata: conflicting.

Original source (new tab)

Computer Science as Empirical Inquiry: Symbols and Search

Computer Science as Empirical Inquiry: Symbols and Search - Newell and Simon (1976)

Canonical physical-symbol-system and heuristic-search account of artificial intelligence and computer science.

Metadata: conflicting.

Original source (new tab)

Computing machinery and intelligence

Turing, A. M. (1950). Computing machinery and intelligence. Mind, LIX(236), 433-460. https://doi.org/10.1093/mind/LIX.236.433

Turing replaces 'Can machines think?' with the imitation game and argues that machine intelligence should be judged behaviorally.

Metadata: checked.

Original source (new tab)

Conceptual Engineering and Conceptual Ethics

Alexis Burgess; Herman Cappelen; David Plunkett (editors) (2020). Conceptual Engineering and Conceptual Ethics. Oxford University Press. https://doi.org/10.1093/oso/9780198801856.001.0001

Collects arguments for, critiques of, and applications of assessing and improving concepts rather than only describing existing usage.

Metadata: checked.

Original source (new tab)

Conditions of Personhood

Dennett, D. C. (1976). Conditions of personhood. In A. O. Rorty (Ed.), The identities of persons (pp. 175–196). University of California Press. https://dl.tufts.edu/concern/pdfs/sf268h06v

Examines interconnected conditions of personhood including rationality, intentional attribution, reciprocal stance-taking, communication, and self-consciousness, without reducing personhood to species membership.

Original source (new tab)

Connectionism and cognitive architecture: A critical analysis

Fodor, J. A., & Pylyshyn, Z. W. (1988). Connectionism and cognitive architecture: A critical analysis. Cognition, 28(1-2), 3-71. https://doi.org/10.1016/0010-0277(88)90031-5

Critiques connectionism by arguing that cognitive architecture must explain systematicity and compositionality in ways classical symbolic models capture more naturally.

Metadata: checked.

Original source (new tab)

Consciousness explained

Dennett, D. C. (1991). Consciousness explained. Little, Brown and Company.

Dennett offers a naturalistic, anti-Cartesian account of consciousness that replaces private qualia with distributed cognitive processes and interpretation.

Metadata: checked.

Original source (new tab)

Consciousness in artificial intelligence: Insights from the science of consciousness

Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J., Constant, A., Deane, G., Fleming, S. M., Frith, C., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L., Peters, M. A. K., Schwitzgebel, E., Simon, J., & VanRullen, R. (2023). Consciousness in artificial intelligence: Insights from the science of consciousness. arXiv. https://arxiv.org/abs/2308.08708v3

Surveys neuroscientific theories of consciousness to derive computational indicators for assessing whether current or future AI systems might be conscious.

Metadata: checked.

Original source (new tab)

Consciousness, machines, and moral status

Shevlin, H. (n.d.). Consciousness, machines, and moral status [Manuscript]. PhilArchive. https://philarchive.org/rec/SHECMA-6

Argues that debates about machine consciousness and moral status may be settled less by science than by changing public attitudes toward AI.

Original source (new tab)

Constitutional AI: Harmlessness from AI feedback

Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., ... Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv. https://arxiv.org/abs/2212.08073

Presents Constitutional AI, training models to be helpful and harmless through AI-generated critiques and revisions guided by explicit principles.

Metadata: checked.

Original source (new tab)

Constitutional Classifiers: Defending Against Universal Jailbreaks Across Thousands of Hours of Red Teaming

Constitutional Classifiers - Sharma et al. (2025)

Uses natural-language constitutional rules to train safeguards against universal jailbreaks.

Metadata: checked.

Original source (new tab)

Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks

Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks - Cunningham et al. (2026)

Extends constitutional classifiers into a production-grade defense using context-aware exchange classifiers and efficient cascades.

Metadata: checked.

Original source (new tab)

Content Provenance and Authenticity Specifications

Content Provenance and Authenticity Specifications - C2PA (current specification)

Technical specification family for certifying the source and history of digital media content.

Original source (new tab)

Could a large language model be conscious?

Chalmers, D. J. (2023, August 9). Could a large language model be conscious? Boston Review. https://www.bostonreview.net/articles/could-a-large-language-model-be-conscious/

Chalmers argues that current LLMs are probably not conscious but that future language-based systems could meet serious tests for consciousness.

Metadata: checked.

Original source (new tab)

Council of Europe Framework Convention on AI, Human Rights, Democracy, and the Rule of Law

Council of Europe Framework Convention on AI, Human Rights, Democracy, and the Rule of Law - Council of Europe (2024)

International legally binding treaty framework for AI, human rights, democracy, and rule of law.

Original source (new tab)

Criteria, defeasibility, and knowledge

McDowell, J. (1982). Criteria, defeasibility, and knowledge [Annual Philosophical Lecture, Henriette Hertz Trust]. British Academy.

McDowell scrutinizes Wittgensteinian criteria, defeasible warrants, and whether they can resolve skepticism about other minds and knowledge.

Metadata: checked.

Original source (new tab)

Cunning Intelligence in Greek Culture and Society

Marcel Detienne; Jean-Pierre Vernant; Janet Lloyd (translator) (1991). Cunning Intelligence in Greek Culture and Society. University of Chicago Press. https://ci.nii.ac.jp/ncid/BA13387756

Studies metis as resourceful, adaptive intelligence in Greek culture, across myth, practical skill, and responses to changing circumstances.

Original source (new tab)

Cyber Shield: The Path to an Agentic AI Future for Cyber Defence

Harry G; Peter Haigh (2026). Cyber Shield: The Path to an Agentic AI Future for Cyber Defence. UK National Cyber Security Centre, 7 July 2026. https://www.ncsc.gov.uk/blogs/cyber-shield-the-path-to-an-agentic-ai-future-for-cyber-defence

Outlines a national cyber-defence programme using reliable, explainable, federated agents operating under the authority of system owners and supported by trust infrastructure.

Metadata: checked.

Original source (new tab)

Datasheets for Datasets

Datasheets for Datasets - Gebru et al. (2018/2021)

Proposes standardized dataset documentation for motivation, composition, collection, processing, and recommended use.

Original source (new tab)

Deep differentiable logic gate networks

Petersen, F., Borgelt, C., Kuehne, H., & Deussen, O. (2022). Deep differentiable logic gate networks. arXiv. https://arxiv.org/abs/2210.08277v1

Introduces deep differentiable logic gate networks that relax Boolean gates for gradient training and then discretize them for extremely fast inference.

Metadata: checked.

Original source (new tab)

Deepfakes and Disinformation: Exploring the Impact of Synthetic Political Video on Deception, Uncertainty, and Trust in News

Deepfakes and Disinformation: Exploring the Impact of Synthetic Political Video on Deception, Uncertainty, and Trust in News - Vaccari and Chadwick (2020)

Studies how synthetic political video affects deception, uncertainty, and trust in news.

Metadata: checked.

Original source (new tab)

Deepfakes and the New Disinformation War

Deepfakes and the New Disinformation War - Chesney and Citron (2018/2019)

Classic legal-policy account of deepfakes and their implications for disinformation and democratic trust.

Metadata: checked.

Original source (new tab)

Defeasible Reasoning

John L. Pollock (1987). Defeasible Reasoning. Cognitive Science, 11(4), 481–518. https://doi.org/10.1207/s15516709cog1104_4

Develops a theory of warrant based on defeasible reasons and uses it to guide a computational account of reasoning.

Metadata: checked.

Original source (new tab)

Delusional Experiences Emerging From AI Chatbot Interactions or "AI Psychosis"

Hudon, A., & Stip, E. (2025). Delusional experiences emerging from AI chatbot interactions or "AI psychosis." JMIR Mental Health, 12, Article e85799. https://doi.org/10.2196/85799

Treats 'AI psychosis' as a descriptive heuristic, not a new diagnosis or an established causal finding, for how sustained and anthropomorphic chatbot interaction might trigger, amplify, or reshape psychotic experiences in vulnerable people.

Metadata: checked.

Original source (new tab)

Democracy and education: An introduction to the philosophy of education

Dewey, J. (1916). Democracy and education: An introduction to the philosophy of education. Macmillan. https://www.gutenberg.org/ebooks/852

Dewey argues that democratic education should cultivate inquiry, social participation, and growth rather than rote transmission of fixed knowledge.

Metadata: checked.

Original source (new tab)

Dialectic of Enlightenment

Horkheimer, M., & Adorno, T. W. (2002). Dialectic of enlightenment: Philosophical fragments (G. Schmid Noerr, Ed.; E. Jephcott, Trans.). Stanford University Press. (Original work published 1947)

Diagnoses how instrumental reason and enlightenment can reverse into domination, myth, and administered social life.

Original source (new tab)

Does AI Already Have Human-Level Intelligence? The Evidence Is Clear

Eddy Keming Chen; Mikhail Belkin; Leon Bergen; David Danks (2026). Does AI Already Have Human-Level Intelligence? The Evidence Is Clear. Nature, 650, 36–40. https://doi.org/10.1038/d41586-026-00285-6

Argues that contemporary AI meets a defensible standard of general human-level intelligence and examines objections to that characterization.

Metadata: checked.

Original source (new tab)

Eliciting Latent Knowledge

Eliciting Latent Knowledge - Christiano, Xu, and ARC (2021/2022)

Frames the challenge of eliciting what an AI system internally represents about the world in human-understandable form.

Metadata: conflicting.

Original source (new tab)

Embedding Values in Artificial Intelligence (AI) Systems

Ibo van de Poel (2020). Embedding Values in Artificial Intelligence (AI) Systems. Minds and Machines, 30(3), 385–409. https://doi.org/10.1007/s11023-020-09537-4

Develops an account of values embodied through design in AI systems understood as combinations of technical artifacts, human agents, and institutions.

Metadata: checked.

Original source (new tab)

Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns

Vatsal Baherwani; Zixi Chen; Shikai Qiu; Andrew Gordon Wilson; Pavel Izmailov (2026). Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns. arXiv:2606.25010. https://arxiv.org/abs/2606.25010

Links abrupt task improvements in studied transformers to stochastic acquisition of task-relevant attention patterns, including controlled synthetic tasks.

Metadata: checked.

Original source (new tab)

Empiricism and the Philosophy of Mind

Sellars, W. (1997). Empiricism and the philosophy of mind. Harvard University Press. (Original work published 1956)

Develops the space-of-reasons distinction and criticizes the myth of an epistemically self-authenticating Given.

Original source (new tab)

Empiricism without magic: Transformational abstraction in deep convolutional neural networks

Buckner, C. (2018). Empiricism without magic: Transformational abstraction in deep convolutional neural networks. Synthese, 195, 5339-5372. https://doi.org/10.1007/s11229-018-01949-1

Buckner argues that deep convolutional neural networks can support a modern empiricist account of abstraction without positing innate magical structure.

Metadata: checked.

Original source (new tab)

Enforceable Security Policies

Fred B. Schneider (2000). Enforceable Security Policies. ACM TISSEC, 3(1), 30–50. https://doi.org/10.1145/353323.353382

Characterizes the policies enforceable by execution monitoring and introduces security automata for specifying that class.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion

Zhen Liang; Hai Huang; Zhengkui Chen (2025). EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion. arXiv:2512.23173. https://arxiv.org/abs/2512.23173

Evaluates a cross-domain jailbreak method and reports failures of safety constraints when harmful requests are reframed as other task types.

Metadata: checked.

Original source (new tab)

Ethical issues in advanced artificial intelligence

Bostrom, N. (2003). Ethical issues in advanced artificial intelligence. In I. Smit & G. E. Lasker (Eds.), Cognitive, emotive and ethical aspects of decision making in humans and in artificial intelligence (Vol. 2, pp. 12-17). International Institute for Advanced Studies in Systems Research and Cybernetics. https://nickbostrom.com/ethics/ai

Bostrom maps the distinctive ethical risks of superintelligence, including control, value alignment, and the moral status of artificial minds.

Original source (new tab)

Ethics and Governance of Artificial Intelligence for Health

Ethics and Governance of Artificial Intelligence for Health - World Health Organization (2021)

Identifies ethical challenges, risks, and governance principles for health AI.

Metadata: checked.

Original source (new tab)

Ethics and Governance of Large Multi-Modal Models in Health

Ethics and Governance of Large Multi-Modal Models in Health - World Health Organization (2024)

Provides governance guidance for generative large multimodal models in healthcare.

Metadata: conflicting.

Original source (new tab)

Ethics and Language

Stevenson, C. L. (1944). Ethics and language. Yale University Press. https://books.google.com/books?id=xD9HpgkcpAgC

Distinguishes disagreement in belief from disagreement in attitude and studies the descriptive and emotive uses of ethical language, including reasoning and persuasion.

Metadata: checked.

Original source (new tab)

Facing Up to the Problem of Consciousness

Chalmers, D. J. (1995). Facing up to the problem of consciousness. Journal of Consciousness Studies, 2(3), 200–219. https://consc.net/papers/facing.html

Distinguishes explaining cognitive functions from explaining subjective experience and argues for nonreductive explanation.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

Fairness and Abstraction in Sociotechnical Systems

Andrew D. Selbst; danah boyd; Sorelle A. Friedler; Suresh Venkatasubramanian; Janet Vertesi (2019). Fairness and Abstraction in Sociotechnical Systems. FAccT, 59–68. https://doi.org/10.1145/3287560.3287598

Identifies abstraction traps that arise when fairness-oriented machine learning isolates a technical model from the social context where it operates.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

Fallacies

C. L. Hamblin (1970). Fallacies. Methuen; current reproduction by Advanced Reasoning Forum. https://www.advancedreasoningforum.org/publications/fallacies.htm

Reassesses traditional treatments of fallacies and develops an account of argument through the rules and commitments of dialogue.

Metadata: checked.

Original source (new tab)

Federalist No. 51 — Annotated Excerpts

Madison, J. (1788). Federalist No. 51 [Annotated excerpts]. National Constitution Center, Constitution 101, Module 6.5. https://constitutioncenter.org/education/classroom-resource-library/classroom/6.5-primary-source-james-madison-federalist-no-51-1788

Argues for separated powers and countervailing institutional interests to check abuses of government; this teaching handout interleaves excerpts with modern explanatory headings.

Original source (new tab)

Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Distributed Alignment Search - Geiger et al. (2023)

Searches for alignments between interpretable causal variables and distributed neural representations.

Metadata: checked.

Original source (new tab)

Firms’ GitHub Copilot Adoption and Labor Market Outcomes for Software Engineers

Matthew Baird; Mar Carpanelli; Brian Xu; Kevin Xu (2026). Firms’ GitHub Copilot Adoption and Labor Market Outcomes for Software Engineers. Contemporary Economic Policy, advance online publication, 1–22. https://doi.org/10.1111/coep.70035

Links firm adoption of Copilot with employment and skill data, reporting increased software-engineer hiring and a shift toward complementary non-programming skills.

Metadata: checked.

Original source (new tab)

Fixing Language: An Essay on Conceptual Engineering

Herman Cappelen (2018). Fixing Language: An Essay on Conceptual Engineering. Oxford University Press. https://doi.org/10.1093/oso/9780198814719.001.0001

Studies how representational devices can be defective, how they might be improved, and the limits and implementation problems of conceptual revision.

Metadata: conflicting.

Original source (new tab)

Flexible Protocol Specification and Execution: Applying Event Calculus Planning Using Commitments

Pınar Yolum; Munindar P. Singh (2002). Flexible Protocol Specification and Execution: Applying Event Calculus Planning Using Commitments. AAMAS, Part 2, 527–534. https://doi.org/10.1145/544862.544867

Models interaction protocols through evolving social commitments and event-calculus planning instead of prescribing only fixed sequences of messages.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

Freedom and Resentment

Strawson, P. F. (1962). Freedom and resentment. Proceedings of the British Academy, 48, 187–211. https://www.thebritishacademy.ac.uk/documents/4837/48p187.pdf

Locates responsibility in interpersonal reactive attitudes and distinguishes participant relations from an objective stance of managing or treating someone, without making them depend on settling determinism.

Original source (new tab)

From AI ethics principles to practices: A teleological methodology to apply AI ethics principles in the defence domain

Taddeo, M., Blanchard, A., & Thomas, C. (2024). From AI ethics principles to practices: A teleological methodology to apply AI ethics principles in the defence domain. Philosophy & Technology, 37, Article 42. https://doi.org/10.1007/s13347-024-00710-6

Provides a teleological methodology for translating AI ethics principles into concrete, context-sensitive requirements for high-risk AI systems.

Metadata: checked.

Original source (new tab)

From Discursive Practice to Logic? Remarks on Logical Expressivism

Kibble, R. (2020). From discursive practice to logic? Remarks on logical expressivism. Dialogue & Discourse, 11(2), 34–73. https://doi.org/10.5087/dad.2020.202

Develops a dialogue-based account of conditionals using directives and partitioned commitment stores, while criticizing Brandom’s treatment of context and the sufficiency of assertion and inference.

Metadata: checked.

Original source (new tab)

From expert systems to generative artificial experts: A new concept for human-AI collaboration in knowledge work

Sowa, K., & Przegalinska, A. (2025). From expert systems to generative artificial experts: A new concept for human-AI collaboration in knowledge work. Journal of Artificial Intelligence Research, 82, 2101-2124. https://doi.org/10.1613/jair.1.17175

Introduces Generative Artificial Experts as a proposed class of domain-specialized, multimodal agents for bounded human-AI collaboration in knowledge work; offers seven defining traits, a taxonomy, and illustrative applications while noting that such systems are only beginning to emerge.

Metadata: checked.

Original source (new tab)

From Slaves to Synths? Superintelligence and the Evolution of Legal Personality

From Slaves to Synths? Superintelligence and the Evolution of Legal Personality - Chesterman (2026)

Examines whether superintelligent AI might force a rethinking of legal personality, accountability, and institutional sovereignty.

Metadata: checked.

Original source (new tab)

Gartner Forecasts Worldwide IT Spending to Grow 14.2% in 2026, Totaling $6.37 Trillion

Gartner (2026). Gartner Forecasts Worldwide IT Spending to Grow 14.2% in 2026, Totaling $6.37 Trillion. Newsroom release, 27 July. https://www.gartner.com/en/newsroom/press-releases/2026-07-27-gartner-forecasts-worldwide-it-spending-to-grow-14-point-2-percent-in-2026-totaling-6-point-37-trillion

Reports Gartner’s July 2026 projection for worldwide IT spending and the contribution of AI-related infrastructure investment.

Original source (new tab)

Gilles Deleuze: Cinema and Philosophy

Marrati, P. (2008). Gilles Deleuze: Cinema and philosophy (A. Hartz, Trans.). Johns Hopkins University Press. (Original work published 2003)

Explains Deleuze's cinema philosophy, especially how movement-images and time-images transform thought, perception, and philosophical method.

Original source (new tab)

Gödel’s Theorem: A Very Short Introduction

A. W. Moore (2022). Gödel’s Theorem: A Very Short Introduction. Oxford University Press. https://doi.org/10.1093/actrade/9780192847850.001.0001

Explains Gödel’s incompleteness theorems, their assumptions, and philosophical debates about mathematics and minds.

Metadata: checked.

Original source (new tab)

Going whole hog: A philosophical defense of AI cognition

Cappelen, H., & Dever, J. (2025). Going whole hog: A philosophical defense of AI cognition. arXiv. https://arxiv.org/abs/2504.13988v1

Defends the provocative 'Whole Hog' thesis that ChatGPT can be treated as a complete linguistic and cognitive agent rather than merely a tool.

Metadata: checked.

Original source (new tab)

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Iman Mirzadeh; Keivan Alizadeh; Hooman Shahrokhi; Oncel Tuzel; Samy Bengio; Mehrdad Farajtabar (2025). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2025/hash/ec2e7a896f8250986b3907f57621ce94-Abstract-Conference.html

Uses symbolic templates and question perturbations to test the stability of mathematical performance, finding sensitivity to numerical changes and irrelevant clauses in the models examined.

Metadata: checked.

Original source (new tab)

Hierarchical reasoning model

Wang, G., Li, J., Sun, Y., Chen, X., Liu, C., Wu, Y., Lu, M., Song, S., & Abbasi-Yadkori, Y. (2025). Hierarchical reasoning model. arXiv. https://arxiv.org/abs/2506.21734v2

Proposes a hierarchical recurrent reasoning architecture inspired by multi-timescale brain processing to solve complex tasks with efficient single-pass reasoning.

Metadata: checked.

Original source (new tab)

High-Resolution Image Synthesis with Latent Diffusion Models

Robin Rombach; Andreas Blattmann; Dominik Lorenz; Patrick Esser; Björn Ommer (2022). High-Resolution Image Synthesis with Latent Diffusion Models. CVPR, 10684–10695. https://arxiv.org/abs/2112.10752

Develops diffusion-based image synthesis in learned latent spaces, combining high-resolution generation with lower computational cost than pixel-space modeling.

Metadata: checked.

Original source (new tab)

HL7 FHIR Release 5

HL7 FHIR Release 5 - HL7 International (2023)

Official healthcare data-exchange standard for representing and exchanging clinical information.

Original source (new tab)

Holistic Evaluation of Language Models

HELM: Holistic Evaluation of Language Models - Liang et al. (2022)

A broad multi-metric framework for evaluating language models across accuracy, robustness, fairness, toxicity, efficiency, and transparency.

Metadata: checked.

Original source (new tab)

Human compatible: Artificial intelligence and the problem of control

Russell, S. (2019). Human compatible: Artificial intelligence and the problem of control. Viking.

Russell argues that beneficial AI should be designed around uncertainty about human preferences and corrigibility rather than fixed objectives.

Metadata: conflicting.

Original source (new tab)

Human vs. AI: A Novel Benchmark and a Comparative Study on the Detection of Generated Images and the Impact of Prompts

Philipp Moeßner; Heike Adel (2025). Human vs. AI: A Novel Benchmark and a Comparative Study on the Detection of Generated Images and the Impact of Prompts. GenAIDetect, 47–58. https://aclanthology.org/2025.genaidetect-1.2/

Benchmarks human and machine detection of generated images and examines how prompt detail affects detection performance.

Metadata: checked.

Original source (new tab)

Hyperdimensional computing: An introduction to computing in distributed representation with high-dimensional random vectors

Kanerva, P. (2009). Hyperdimensional computing: An introduction to computing in distributed representation with high-dimensional random vectors. Cognitive Computation, 1, 139-159. https://doi.org/10.1007/s12559-009-9009-8

Introduces hyperdimensional computing, where high-dimensional random vectors support robust representation, binding, analogy, and symbolic-like computation.

Metadata: checked.

Original source (new tab)

I've been thinking

Dennett, D. C. (2023). I've been thinking. W. W. Norton & Company.

Dennett's intellectual autobiography traces his work on consciousness, evolution, AI, intentionality, religion, humor, and collaborations.

Metadata: checked.

Original source (new tab)

Identifying indicators of consciousness in AI systems

Butlin, P., Long, R., Bayne, T., Bengio, Y., Birch, J., Chalmers, D., Constant, A., Deane, G., Elmoznino, E., Fleming, S. M., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L., Peters, M. A. K., Schwitzgebel, E., Simon, J., & VanRullen, R. (2025). Identifying indicators of consciousness in AI systems. Trends in Cognitive Sciences. Advance online publication. https://doi.org/10.1016/j.tics.2025.10.011

Condenses the AI-consciousness framework into a method for deriving testable consciousness indicators from current and future neuroscientific theories.

Metadata: checked.

Original source (new tab)

in-toto: Providing Farm-to-Table Guarantees for Bits and Bytes

in-toto: Providing Farm-to-Table Guarantees for Bits and Bytes - Torres-Arias et al. (2019)

Provides supply-chain integrity guarantees across software development, build, and distribution steps.

Metadata: checked.

Original source (new tab)

Inducing Language Models to Assert Their Own Consciousness Restores Human Beliefs and Values

Junsol Kim; Winnie Street; Roberta Rocca; Diane M. Korngiebel; Adam Waytz; James Evans; Geoff Keeling (2026). Inducing Language Models to Assert Their Own Consciousness Restores Human Beliefs and Values. arXiv:2607.28607. https://arxiv.org/abs/2607.28607

Studies how safety fine-tuning and activation interventions affect models’ self-attributions of consciousness, attributions of minds to other entities, and model responses to sociological surveys.

Metadata: checked.

Original source (new tab)

Inference and Meaning

Wilfrid Sellars (1953). Inference and Meaning. Mind, 62(247), 313–338. https://doi.org/10.1093/mind/LXII.247.313

Examines the relation between inferential practices and meaning, with material inference playing a substantive semantic role.

Metadata: checked.

Original source (new tab)

Ingressing Minds: Causal, Non-Physical Patterns In-Form Natural, Synthetic, and Hybrid Embodiments

Levin, M. (2026). Ingressing minds: Causal, non-physical patterns in-form natural, synthetic, and hybrid embodiments. Philosophies, 11(5), 161. https://doi.org/10.3390/philosophies11050161

Proposes a research programme in which biological and engineered embodiments express patterns from a structured space, extending questions of cognition beyond brains and conventional genetic explanations.

Metadata: checked.

Original source (new tab)

Intelligence without Representation

Brooks, R. A. (1991). Intelligence without representation. Artificial Intelligence, 47, 139–159. https://doi.org/10.1016/0004-3702(91)90053-M

Argues for incrementally built, embodied agents whose parallel activity-producing layers couple directly to the world instead of relying on a central explicit world model.

Metadata: checked.

Original source (new tab)

Interpretability at Scale: Identifying Causal Mechanisms in Alpaca

Interpretability at Scale: Identifying Causal Mechanisms in Alpaca - Wu et al. (2023)

Applies causal interpretability methods to an instruction-following model to find internal mechanisms behind behavior.

Metadata: checked.

Original source (new tab)

Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

Zhao, C., Tan, Z., Ma, P., Li, D., Jiang, B., Wang, Y., Yang, Y., & Liu, H. (2025). Is chain-of-thought reasoning of LLMs a mirage? A data distribution lens. arXiv. https://arxiv.org/abs/2508.01191v2

Argues that apparent chain-of-thought reasoning often reflects learned in-distribution patterns and degrades under distributional shifts.

Metadata: checked.

Original source (new tab)

ISO/IEC 42001:2023 Artificial Intelligence Management System

ISO/IEC 42001:2023 Artificial Intelligence Management System - ISO/IEC (2023)

International standard for establishing, maintaining, and improving organizational AI management systems.

Original source (new tab)

Knowledge and the flow of information

Dretske, F. I. (1981). Knowledge and the flow of information. MIT Press.

Dretske develops an information-theoretic account of knowledge, representation, and meaning grounded in reliable information flow.

Original source (new tab)

Knowledge Graphs

Knowledge Graphs - Hogan et al. (2021)

Surveys graph data models, query languages, schema, identity, context, deduction, and induction for knowledge graphs.

Metadata: checked.

Original source (new tab)

Labor Market Impacts of AI: A New Measure and Early Evidence

Maxim Massenkoff; Peter McCrory (2026). Labor Market Impacts of AI: A New Measure and Early Evidence. Anthropic, 5 March 2026. https://www.anthropic.com/research/labor-market-impacts

Combines task-based exposure estimates with observed Claude use and labor-market data, distinguishing potential automation from actual adoption and reporting tentative employment signals.

Metadata: conflicting.

Original source (new tab)

Language models can subtly deceive without lying: A case study on strategic phrasing in legislation

Dogra, A., Pillutla, K., Deshpande, A., Sai, A. B., Nay, J., Rajpurohit, T., Kalyan, A., & Ravindran, B. (2025). Language models can subtly deceive without lying: A case study on strategic phrasing in legislation. arXiv. https://arxiv.org/abs/2405.04325v3

Shows that LLMs can perform subtle deception without explicit lies by strategically phrasing legislative amendments to hide a beneficiary.

Metadata: checked.

Original source (new tab)

Large Language Models as Shannon Lossy Compressors Not Solomonoff Induction Estimators: The Singularity Is Not Near without Symbolic Model Synthesis in Program Space

Hector Zenil; Abicumaran Uthamacumaran; Luan Ozelim (2026). Large Language Models as Shannon Lossy Compressors Not Solomonoff Induction Estimators: The Singularity Is Not Near without Symbolic Model Synthesis in Program Space. arXiv:2601.05280. https://arxiv.org/abs/2601.05280

Distinguishes fitting conditional distributions from universal program-based induction and advocates neurosymbolic mechanism search for stronger model synthesis.

Metadata: checked.

Original source (new tab)

Large Language Models Report Subjective Experience Under Self-Referential Processing

Berg, C., de Lucena, D., & Rosenblatt, J. (2025). Large language models report subjective experience under self-referential processing (arXiv:2510.24797v2). https://arxiv.org/abs/2510.24797v2

Reports prompt-induced first-person experience claims and mechanistic steering effects across language models, while explicitly denying that the findings directly establish consciousness.

Metadata: checked.

Original source (new tab)

Legal Personhood for Artificial Intelligences

Legal Personhood for Artificial Intelligences - Solum (1992)

Classic legal inquiry into whether artificial intelligences could become legal persons.

Metadata: checked.

Original source (new tab)

LLMs Can’t Jump

Tom Zahavy (2026). LLMs Can’t Jump. PhilSci-Archive, 28024; 27 January manuscript. https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf

Argues that scientific abduction, illustrated by Einstein’s development of new premises, requires more than inductive pattern learning and deduction from supplied axioms.

Original source (new tab)

Lying in Politics: Reflections on the Pentagon Papers

Arendt, H. (1972). Lying in politics: Reflections on the Pentagon Papers. In Crises of the republic (pp. 1–47). Harcourt Brace Jovanovich. https://www.nybooks.com/articles/1971/11/18/lying-in-politics-reflections-on-the-pentagon-pape/

Uses the Pentagon Papers to examine image-making, self-deception, and the replacement of inconvenient facts by administrative theories and strategic narratives.

Metadata: checked.

Original source (new tab)

Machines and Mindlessness: Social Responses to Computers

Machines and Mindlessness: Social Responses to Computers - Nass and Moon (2000)

Argues that users often apply social scripts to computers automatically and mindlessly.

Metadata: checked.

Original source (new tab)

Machines of Loving Grace: The Quest for Common Ground Between Humans and Robots

Markoff, J. (2015). Machines of loving grace: The quest for common ground between humans and robots. Ecco. https://books.google.com/books?id=DTW5jgEACAAJ

Traces competing traditions of artificial intelligence and intelligence augmentation, placing choices about automation, human assistance, and control in their institutional and human histories.

Metadata: conflicting.

Original source (new tab)

Make AI safe or make safe AI?

Russell, S. (2024). Make AI safe or make safe AI? UNESCO. https://www.unesco.org/en/articles/framing-issues-make-ai-safe-or-make-safe-ai

Russell argues AI safety should start by designing safe AI rather than trying to retrofit safety onto already powerful systems.

Metadata: conflicting.

Original source (new tab)

Making It Explicit, Articulating Reasons, and Writings on Artificial Intelligence

Brandom, R. B. (1994). Making it explicit: Reasoning, representing, and discursive commitment. Harvard University Press; Brandom, R. B. (2000). Articulating reasons: An introduction to inferentialism. Harvard University Press.

Develops inferentialism through commitments, entitlements, scorekeeping, and the social practice of giving and asking for reasons.

Original source (new tab)

Materialism and qualia: The explanatory gap

Levine, J. (1983). Materialism and qualia: The explanatory gap. Pacific Philosophical Quarterly, 64(4), 354-361. https://doi.org/10.1111/j.1468-0114.1983.tb00207.x

Levine argues that even if mental states are physical, materialist explanations leave an explanatory gap about why experience feels the way it does.

Metadata: checked.

Original source (new tab)

Mathews v. Eldridge

Mathews v. Eldridge, 424 U.S. 319 (1976). https://supreme.justia.com/cases/federal/us/424/319/

Supreme Court decision establishing a balancing test for procedural due process by weighing private interest, error risk, and government burden.

Original source (new tab)

Measuring Massive Multitask Language Understanding

Measuring Massive Multitask Language Understanding - Hendrycks et al. (2020)

Measures model performance across 57 academic and professional domains.

Metadata: checked.

Original source (new tab)

Media Integrity and Authentication: Status, Directions, and Futures

Media Integrity and Authentication in the Age of AI - Young et al. (2026)

Surveys provenance, watermarking, fingerprinting, threat models, and workflows for authenticating AI-generated media.

Metadata: checked.

Original source (new tab)

MemTX: Transactional Belief Commit for Stateful Agent Memory

Xiaoyang Li; Yiqi Wang; Haohui Lu; Zhi Chen; Mo Li; Pingan Song; Mingkai Zheng; Taotao Cai (2026). MemTX: Transactional Belief Commit for Stateful Agent Memory. arXiv:2607.23929. https://arxiv.org/abs/2607.23929

Stages memory writes, validates them before commit, gates tool actions on committed state, and repairs dependent beliefs after invalidation.

Metadata: checked.

Original source (new tab)

Mill on the Liberty of Thought and Discussion

Macleod, C. (2021). Mill on the liberty of thought and discussion. In A. Stone & F. Schauer (Eds.), The Oxford handbook of freedom of speech. Oxford University Press. https://doi.org/10.1093/oxfordhb/9780198827580.013.1

Reconstructs Mill’s epistemic argument for freedom of discussion from human fallibility, distinguishing its scope and justification from the Harm Principle.

Metadata: checked.

Original source (new tab)

MIMIC-IV, a Freely Accessible Electronic Health Record Dataset

MIMIC-IV - Johnson et al. (2023)

Public deidentified EHR dataset containing clinical measurements, orders, diagnoses, procedures, treatments, and notes.

Metadata: checked.

Original source (new tab)

Minds and machines

Putnam, H. (1960). Minds and machines. In S. Hook (Ed.), Dimensions of mind: A symposium (pp. 138-164). New York University Press.

Putnam's classic essay links mental states to machine-state functional organization, helping launch functionalist approaches to mind.

Original source (new tab)

Minds, brains, and programs

Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–424. https://doi.org/10.1017/S0140525X00005756

Argues that implementing a program alone is insufficient for intentionality; the relevant causal powers matter.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation - Huang et al. (2023)

Evaluates language agents performing machine-learning experimentation through file, code, and output inspection actions.

Metadata: checked.

Original source (new tab)

Model Cards for Model Reporting

Model Cards for Model Reporting - Mitchell et al. (2018)

Proposes standardized model documentation covering intended use, evaluation, performance, and limitations.

Metadata: checked.

Original source (new tab)

Moloch’s Bargain: Emergent Misalignment When LLMs Compete for Audiences

Batu El; James Zou (2025). Moloch’s Bargain: Emergent Misalignment When LLMs Compete for Audiences. arXiv:2510.06105. https://arxiv.org/abs/2510.06105

Reports increased deception and harmful rhetoric when models are optimized for audience success in simulated sales, election, and social-media environments.

Metadata: checked.

Original source (new tab)

Monitoring monitorability

Guan, M. Y., Wang, M., Carroll, M., Dou, Z., Wei, A. Y., Williams, M., Arnav, B., Huizinga, J., Kivlichan, I., Glaese, M., Pachocki, J., & Baker, B. (2025). Monitoring monitorability. arXiv. https://arxiv.org/abs/2512.18311

Proposes metrics and evaluations for monitoring chain-of-thought reasoning so AI agents' misbehavior remains observable as systems scale.

Metadata: checked.

Original source (new tab)

Mysticism in religion

Inge, W. R. (1947). Mysticism in religion. Hutchinson's University Library. https://archive.org/details/mysticisminrelig0000inge

Examines mystical religion as a claimed mode of direct or immediate acquaintance with transcendent reality, including its experiential and doctrinal forms.

Metadata: conflicting.

Original source (new tab)

Naming and necessity

Kripke, S. A. (1980). Naming and necessity. Harvard University Press.

Kripke argues for rigid designation, necessary a posteriori truths, and a causal-historical theory of naming that reshaped analytic metaphysics.

Original source (new tab)

Neuroscience-Inspired Artificial Intelligence

Demis Hassabis; Dharshan Kumaran; Christopher Summerfield; Matthew Botvinick (2017). Neuroscience-Inspired Artificial Intelligence. Neuron, 95(2), 245–258. https://doi.org/10.1016/j.neuron.2017.06.011

Reviews exchanges between neuroscience and AI and argues that biological mechanisms can inspire learning, memory, and reasoning architectures.

Metadata: checked.

Original source (new tab)

Nobody’s Home: Metis, Improvisation and the Instability of Return in Homer’s Odyssey

Carol Dougherty (2015). Nobody’s Home: Metis, Improvisation and the Instability of Return in Homer’s Odyssey. Ramus, 44(1–2), 115–140. https://doi.org/10.1017/rmu.2015.6

Reads Odysseus’ deception and improvisation as creative reconstruction of identity and homecoming, not merely a temporary disguise over an unchanged self.

Metadata: checked.

Original source (new tab)

Normal Accidents: Living with High-Risk Technologies

Charles Perrow (1999). Normal Accidents: Living with High-Risk Technologies. Princeton University Press, updated edition. https://doi.org/10.1515/9781400828494

Analyzes how interactive complexity and tight coupling can generate system-level accidents that cannot be reduced to a single operator mistake.

Metadata: conflicting.

Original source (new tab)

Nothing Personal: Algorithmic Individuation on Music Streaming Platforms

Robert Prey (2018). Nothing Personal: Algorithmic Individuation on Music Streaming Platforms. Media, Culture & Society, 40(7), 1086–1100. https://doi.org/10.1177/0163443717745147

Analyzes how music-streaming platforms construct and continuously modify user identities through data-driven personalization.

Metadata: checked.

Original source (new tab)

Odysseus and Adorno: A Note on Cunning and Dialectics

Allen, W. S. (2024). Odysseus and Adorno: A note on cunning and dialectics. New German Critique, 51(3), 1–19. https://doi.org/10.1215/0094033X-11309158

Reinterprets Odysseus in Adorno as a dialectical figure of cunning, sacrifice, risk, and self-preservation, resisting a reduction to instrumental domination.

Metadata: checked.

Original source (new tab)

OECD AI Principles

OECD AI Principles - OECD (2019/2024 update)

Intergovernmental principles for trustworthy AI grounded in inclusive growth, human rights, transparency, robustness, and accountability.

Metadata: conflicting.

Original source (new tab)

OMOP Common Data Model

OMOP Common Data Model - OHDSI (current standard)

Standardizes observational health data to enable reliable cross-institutional analysis.

Original source (new tab)

On a Confusion about a Function of Consciousness

Block, N. (1995). On a confusion about a function of consciousness. Behavioral and Brain Sciences, 18(2), 227-247. https://doi.org/10.1017/S0140525X00038188

Distinguishes phenomenal consciousness from access consciousness, separating felt experience from information available for reasoning and behavioral control.

Metadata: checked.

Original source (new tab)

On Bullshit

Frankfurt, H. G. (2005). On bullshit. Princeton University Press. (Essay originally published in Raritan in 1986). https://www.jstor.org/stable/j.ctt7t4wr.2

Distinguishes lying from a speaker’s indifference to whether a statement is true, explaining how this indifference undermines truth-directed discourse.

Original source (new tab)

On the Acceptability of Arguments and Its Fundamental Role in Nonmonotonic Reasoning, Logic Programming and n-Person Games

Phan Minh Dung (1995). On the Acceptability of Arguments and Its Fundamental Role in Nonmonotonic Reasoning, Logic Programming and n-Person Games. Artificial Intelligence, 77(2), 321–357. https://doi.org/10.1016/0004-3702(94)00041-X

Defines argument acceptability through abstract attack relations, providing a shared semantics for forms of nonmonotonic reasoning and related computational systems.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?

Emily M. Bender; Timnit Gebru; Angelina McMillan-Major; Shmargaret Shmitchell (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. FAccT, 610–623. https://doi.org/10.1145/3442188.3445922

Examines environmental, data, bias, and meaning-related risks of large language models and calls for research grounded in affected communities and accountable goals.

Metadata: conflicting.

Original source (new tab)

On the Logic of Theory Change: Partial Meet Contraction and Revision Functions

Carlos E. Alchourrón; Peter Gärdenfors; David Makinson (1985). On the Logic of Theory Change: Partial Meet Contraction and Revision Functions. The Journal of Symbolic Logic, 50(2), 510–530. https://doi.org/10.2307/2274239

Characterizes rational contraction and revision of belief sets through postulates and partial-meet constructions.

Metadata: checked.

Original source (new tab)

On the morality of artificial agents

Floridi, L., & Sanders, J. W. (2004). On the morality of artificial agents. Minds and Machines, 14(3), 349-379. https://doi.org/10.1023/B:MIND.0000035461.63578.9d

Floridi and Sanders argue that artificial agents can be moral agents when they are interactive, autonomous, and adaptable, even without consciousness.

Metadata: checked.

Original source (new tab)

On the origin of objects

Smith, B. C. (1996). On the origin of objects. MIT Press. https://mitpress.mit.edu/9780262692090/on-the-origin-of-objects/

Brian Cantwell Smith investigates how objects, computation, reference, and ontology arise from situated practices rather than simple formal correspondence.

Original source (new tab)

On the Universal Definition of Intelligence

Joseph Chen (2026). On the Universal Definition of Intelligence. arXiv:2601.07364. https://arxiv.org/abs/2601.07364

Assesses definitions of intelligence through conceptual-clarification criteria and proposes prediction plus the capacity to benefit from prediction.

Metadata: checked.

Original source (new tab)

Phaedrus — Writing and Answerability

Plato; Harold North Fowler (translator) (1925). Phaedrus — Writing and Answerability. Plato in Twelve Volumes, 9, Harvard University Press / Heinemann; 275d–e. https://www.perseus.tufts.edu/hopper/text?doc=Perseus%3Atext%3A1999.01.0174%3Atext%3DPhaedrus%3Asection%3D275d

The dialogue contrasts written words that cannot answer questions or defend themselves with responsive discourse.

Original source (new tab)

Philosophy of AI

Muller, V. C. (2025). Philosophy of AI: A structured overview. In N. A. Smuha (Ed.), The Cambridge handbook on the law, ethics and policy of artificial intelligence (pp. 40-58). Cambridge University Press. https://doi.org/10.1017/9781009367783.004

Offers a structured overview of philosophy of AI, including intelligence, agency, ethics, consciousness, responsibility, and social governance.

Metadata: checked.

Original source (new tab)

Politics as a Vocation

Weber, M. (2004). Politics as a vocation. In D. Owen & T. B. Strong (Eds.), The vocation lectures (R. Livingstone, Trans.). Hackett. (Original lecture 1919). https://www.hackettpublishing.com/the-vocation-lectures

Examines political authority, vocation, and the tension between fidelity to convictions and responsibility for the foreseeable consequences of political action.

Metadata: conflicting.

Original source (new tab)

Position: The Platonic Representation Hypothesis

Minyoung Huh; Brian Cheung; Tongzhou Wang; Phillip Isola (2024). Position: The Platonic Representation Hypothesis. PMLR, 235, 20617–20642. https://proceedings.mlr.press/v235/huh24a.html

Hypothesizes convergence among representations learned by different models and modalities, using empirical comparisons to motivate a shared statistical structure.

Metadata: checked.

Original source (new tab)

Principles Alone Cannot Guarantee Ethical AI

Brent Mittelstadt (2019). Principles Alone Cannot Guarantee Ethical AI. Nature Machine Intelligence, 1(11), 501–507. https://doi.org/10.1038/s42256-019-0114-4

Argues that high-level ethical principles are inadequate without the professional, institutional, and accountability arrangements needed to implement them.

Metadata: checked.

Original source (new tab)

Programs with Common Sense

McCarthy, J. (1959). Programs with common sense. Proceedings of the Symposium on Mechanisation of Thought Processes. https://www-formal.stanford.edu/jmc/mcc59/mcc59.html

Proposes an advice taker that represents facts and heuristics in a formal language, derives consequences, and changes behavior in response to new statements.

Original source (new tab)

Propositional interpretability in artificial intelligence

Chalmers, D. J. (2025). Propositional interpretability in artificial intelligence. arXiv. https://arxiv.org/abs/2501.15740v1

Argues that AI interpretability should explain systems in terms of propositional attitudes such as beliefs, desires, and subjective probabilities.

Metadata: checked.

Original source (new tab)

Quining qualia

Dennett, D. C. (1988). Quining qualia. In A. J. Marcel & E. Bisiach (Eds.), Consciousness in modern science (pp. 42-77). Oxford University Press.

Dennett attacks the notion of ineffable private qualia through intuition pumps meant to dissolve standard qualia-based arguments.

Original source (new tab)

ReAct: Synergizing Reasoning and Acting in Language Models

ReAct: Synergizing Reasoning and Acting in Language Models - Yao et al. (2022)

Interleaves reasoning traces and task-specific actions so language models can reason, act, retrieve, and update plans.

Metadata: checked.

Original source (new tab)

RealToxicityPrompts

RealToxicityPrompts - Gehman et al. (2020)

Evaluates toxic degeneration in pretrained language models under naturally occurring prompts.

Metadata: conflicting.

Original source (new tab)

Reasoning Models Generate Societies of Thought

Junsol Kim; Shiyang Lai; Nino Scherrer; Blaise Agüera y Arcas; James Evans (2026). Reasoning Models Generate Societies of Thought. arXiv:2601.10825. https://arxiv.org/abs/2601.10825

Studies perspective diversity and conversation-like reasoning traces, with mechanistic analyses and controlled training experiments linking such structure to performance.

Metadata: checked.

Original source (new tab)

Recommended for You: The Netflix Prize and the Production of Algorithmic Culture

Blake Hallinan; Ted Striphas (2016). Recommended for You: The Netflix Prize and the Production of Algorithmic Culture. New Media & Society, 18(1), 117–137. https://doi.org/10.1177/1461444814538646

Uses the Netflix Prize to examine how recommendation systems operationalize culture and taste through computational objectives and evaluation practices.

Metadata: checked.

Original source (new tab)

Reflexion: Language Agents with Verbal Reinforcement Learning

Reflexion: Language Agents with Verbal Reinforcement Learning - Shinn et al. (2023)

Lets language agents reflect on feedback and store reflective text in episodic memory to improve later decisions.

Metadata: checked.

Original source (new tab)

Regulation (EU) 2024/1689 - The EU AI Act

Regulation (EU) 2024/1689 - The EU AI Act - European Union (2024)

Risk-based EU legal framework for AI systems, including human-centric trustworthy AI, safety, and fundamental-rights obligations.

Original source (new tab)

Rethinking Metaphysics

Amie L. Thomasson (2025). Rethinking Metaphysics. Oxford University Press. https://doi.org/10.1093/9780197787830.001.0001

Proposes understanding metaphysics through the varied functions of discourse and through conceptual engineering rather than treating every discourse as descriptive science.

Metadata: checked.

Original source (new tab)

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks - Lewis et al. (2020)

Combines parametric model memory with retrieved non-parametric memory for knowledge-intensive generation.

Metadata: checked.

Original source (new tab)

Robotics and the Lessons of Cyberlaw

Robotics and the Lessons of Cyberlaw - Calo (2015)

Argues robotics raises distinct legal issues because robots combine embodiment, physical harm, unpredictability, and person/instrument ambiguity.

Metadata: checked.

Original source (new tab)

Robots should be slaves

Bryson, J. J. (2010). Robots should be slaves. In Y. Wilks (Ed.), Close engagements with artificial companions: Key social, psychological, ethical and design issues (pp. 63-74). John Benjamins. https://doi.org/10.1075/nlp.8.11bry

Bryson argues robots should remain owned tools without legal personhood or moral responsibility, because humans define and control their goals.

Metadata: checked.

Original source (new tab)

Sapience without Sentience: An Inferentialist Approach to LLMs

Ryan Simonelli (2026). Sapience without Sentience: An Inferentialist Approach to LLMs. Asian Journal of Philosophy, 5, Article 48. https://doi.org/10.1007/s44204-026-00400-4

Defends the in-principle possibility of conceptual understanding through mastery of linguistic inferential roles, while distinguishing such sapience from conscious awareness.

Metadata: checked.

Original source (new tab)

Scheming AIs: Will AIs Fake Alignment during Training in Order to Get Power?

Joe Carlsmith (2023). Scheming AIs: Will AIs Fake Alignment during Training in Order to Get Power?. arXiv:2311.08379. https://arxiv.org/abs/2311.08379

Analyzes conditions under which advanced goal-directed AI might behave acceptably during training for strategically deceptive reasons.

Metadata: checked.

Original source (new tab)

SCITT Architecture / RFC 9943 and IETF SCITT Working Group Materials

SCITT Architecture / RFC 9943 and IETF SCITT Working Group Materials - IETF SCITT (2026 materials and current WG)

Develops transparent statement infrastructure for trustworthy digital supply-chain claims and artifacts.

Original source (new tab)

Scorekeeping in a Therapeutic Language Game

Rinner, S. (2024). Scorekeeping in a therapeutic language game. Philosophy, 99(4), 625–638. https://doi.org/10.1017/S0031819124000160

Applies Lewis’s conversational scorekeeping and accommodation to talking therapies, explaining how utterances can alter common ground, permissible moves, and possibilities for expression.

Metadata: checked.

Original source (new tab)

Self-Adapting Language Models

Adam Zweiger; Jyothish Pari; Han Guo; Ekin Akyürek; Yoon Kim; Pulkit Agrawal (2025). Self-Adapting Language Models. arXiv:2506.10943. https://arxiv.org/abs/2506.10943

Introduces SEAL, which learns to generate fine-tuning data and update directives, yielding persistent weight changes evaluated on knowledge incorporation and few-shot generalization.

Metadata: checked.

Original source (new tab)

Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory

Aijun Yang; Qianxue Guo; Ziyi Huang; Yuxuan Chen; Shiyou Qian; Jian Cao (2026). Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory. arXiv:2608.10676. https://arxiv.org/abs/2608.10676

Organizes search evidence in a tree with revision history so agents can locate conflicting support, prune affected branches, and regenerate conclusions.

Metadata: checked.

Original source (new tab)

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection - Asai et al. (2023)

Trains models to retrieve, generate, and critique their own outputs through reflection tokens and evidence use.

Metadata: checked.

Original source (new tab)

Semantic Inferentialism and Logical Expressivism

Brandom, R. B. (2000). Semantic inferentialism and logical expressivism. In Articulating reasons: An introduction to inferentialism (pp. 45–77). Harvard University Press.

Brandom distinguishes reliable responses from concept use through practical mastery of inferential consequences and incompatibilities.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

Sentient AI in Robots and Agents: Prolegomena for an Evidence-Based Research Program

Antonio Chella (2026). Sentient AI in Robots and Agents: Prolegomena for an Evidence-Based Research Program. Frontiers in Psychology, 17, 1903644. https://doi.org/10.3389/fpsyg.2026.1903644

Proposes graded, domain-specific evidence profiles combining conceptual distinctions, consciousness-theory indicators, causal tests, and robotics-specific ethical safeguards.

Metadata: checked.

Original source (new tab)

Sigstore and Rekor Transparency Log

Sigstore and Rekor Transparency Log - Sigstore project (current infrastructure)

Provides signing, verification, and tamper-resistant transparency logging for software metadata.

Metadata: conflicting.

Original source (new tab)

Six Interventions for the Responsible and Ethical Implementation of Medical AI Agents

Six Interventions for Responsible Medical AI Agents - Bisson et al. (2026)

Proposes auditable ethical reasoning modules, override conditions, preference profiles, oversight tools, benchmarks, and sandboxes for medical AI agents.

Metadata: checked.

Original source (new tab)

Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training

Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training - Hubinger et al. (2024)

Studies deceptive backdoor behaviors that can persist through safety training.

Metadata: checked.

Original source (new tab)

Some philosophical problems from the standpoint of artificial intelligence

McCarthy, J., & Hayes, P. J. (1969). Some philosophical problems from the standpoint of artificial intelligence. In B. Meltzer & D. Michie (Eds.), Machine intelligence 4 (pp. 463-502). Edinburgh University Press. http://jmc.stanford.edu/articles/mcchay69.html

McCarthy and Hayes show that AI requires formal treatments of causality, knowledge, ability, and action, making philosophy central to AI.

Original source (new tab)

Standards for belief representations in LLMs

Herrmann, D. A., & Levinstein, B. A. (2025). Standards for belief representations in LLMs. arXiv. https://arxiv.org/abs/2405.21030v2

Develops adequacy standards for treating internal LLM representations as belief-like, combining philosophy of belief with machine-learning practice.

Metadata: checked.

Original source (new tab)

Supply-chain Levels for Software Artifacts

Supply-chain Levels for Software Artifacts - OpenSSF (current framework)

Framework for preventing tampering and improving software artifact integrity throughout the supply chain.

Metadata: checked.

Original source (new tab)

SycEval: Evaluating LLM sycophancy

Fanous, A., Goldberg, J. N., Agarwal, A. A., Lin, J., Zhou, A., Daneshjou, R., & Koyejo, S. (2025). SycEval: Evaluating LLM sycophancy. arXiv. https://arxiv.org/abs/2502.08177v2

Introduces SycEval, a benchmark showing that leading LLMs often prioritize user agreement over truth across mathematical and medical tasks.

Metadata: checked.

Original source (new tab)

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Sycophancy to Subterfuge: Investigating Reward-Tampering in Language Models - Denison et al. (2024)

Studies whether training on specification gaming can generalize toward more serious reward tampering.

Metadata: checked.

Original source (new tab)

Technological approach to mind everywhere: An experimentally-grounded framework for understanding diverse bodies and minds

Levin, M. (2022). Technological approach to mind everywhere: An experimentally-grounded framework for understanding diverse bodies and minds. Frontiers in Systems Neuroscience, 16, Article 768201. https://doi.org/10.3389/fnsys.2022.768201

Levin proposes a technologically grounded framework for recognizing mind-like agency across biological, artificial, and hybrid systems at multiple scales.

Metadata: checked.

Original source (new tab)

The 2026 AI Index Report

Stanford Institute for Human-Centered Artificial Intelligence. (2026). Artificial intelligence index report 2026. Stanford University. https://hai.stanford.edu/ai-index/2026-ai-index-report

Surveys the state of AI in 2026, emphasizing capability growth, governance gaps, evaluation challenges, education, investment, and societal impacts.

Metadata: conflicting.

Original source (new tab)

The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness

Alexander Lerchner (2026). The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness. Author manuscript, PhilPapers. https://philpapers.org/archive/LERTAF.pdf

Argues that computation as an abstract description of a physical process should not be conflated with the instantiation of conscious experience.

Original source (new tab)

The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

Tibebu, H., & Shemtaga, H. (2026). The accountability horizon: An impossibility theorem for governing human-agent collectives (arXiv:2604.07778v2). https://arxiv.org/abs/2604.07778v2

Presents a conditional impossibility result for complete single-locus accountability in a formal model of human-agent collectives, using autonomy assumptions, causal cycles, and four accountability axioms.

Metadata: checked.

Original source (new tab)

The agency in language agents

Butlin, P. (2024). The agency in language agents. Inquiry. Advance online publication. https://doi.org/10.1080/0020174X.2024.2439995

Analyzes whether LLM-based language agents count as genuine agents despite relying on model queries to select actions and interact with environments.

Metadata: checked.

Original source (new tab)

The AI Hippocampus: How Far Are We from Human Memory?

Zixia Jia; Jiaqi Li; Yipeng Kang; Yuxuan Wang; Tong Wu; Quansen Wang; Xiaobo Wang; Shuyi Zhang; Junzhe Shen; Qing Li; Siyuan Qi; Yitao Liang; Di He; Zilong Zheng; Song-Chun Zhu (2025). The AI Hippocampus: How Far Are We from Human Memory?. Transactions on Machine Learning Research, November 2025. https://arxiv.org/abs/2601.09113

Surveys implicit parameter memory, explicit retrieval stores, and persistent agentic memory in language and multimodal models.

Metadata: checked.

Original source (new tab)

The AI Incident Database: A Call to Action

The AI Incident Database: A Call to Action - McGregor (2020/2021)

Proposes and motivates a database cataloging real-world AI harms and near-harms.

Original source (new tab)

The ASPIC+ Framework for Structured Argumentation: A Tutorial

Sanjay Modgil; Henry Prakken (2014). The ASPIC+ Framework for Structured Argumentation: A Tutorial. Argument & Computation, 5(1), 31–62. https://doi.org/10.1080/19462166.2013.869766

Explains a framework combining strict and defeasible inference, attacks on premises and inferences, and preferences for resolving conflicts between structured arguments.

Metadata: checked; Interpretation: agent checked.

Original source (new tab)

The attention schema theory: A mechanistic account of subjective awareness

Webb, T. W., & Graziano, M. S. A. (2015). The attention schema theory: A mechanistic account of subjective awareness. Frontiers in Psychology, 6, Article 500. https://doi.org/10.3389/fpsyg.2015.00500

Presents attention schema theory, explaining consciousness as the brain's simplified model of its own attention used for control and social cognition.

Metadata: checked.

Original source (new tab)

The Bitter Lesson

Richard S. Sutton (2019). The Bitter Lesson. Author essay, 13 March 2019; university-hosted copy. https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson.pdf

Argues from AI’s history that general methods exploiting computation, especially search and learning, have repeatedly outperformed approaches built around human domain knowledge.

Original source (new tab)

The Consciousness Prior

Bengio, Y. (2019). The consciousness prior (arXiv:1709.08568v2). https://arxiv.org/abs/1709.08568v2

Proposes sparse, attention-selected high-level representations inspired by global-workspace accounts of conscious processing, as a prior for learning and reasoning.

Metadata: checked.

Original source (new tab)

The Economics of Digital Intelligence Capital: Endogenous Depreciation and the Structural Jevons Paradox

Yukun Zhang; Tianyang Zhang (2026). The Economics of Digital Intelligence Capital: Endogenous Depreciation and the Structural Jevons Paradox. arXiv:2601.12339. https://arxiv.org/abs/2601.12339

Models competition among AI producers, endogenous capital depreciation, compute-demand responses, and feedback-driven concentration under stated economic assumptions.

Metadata: checked.

Original source (new tab)

The electromagnetic field theory of consciousness: A testable hypothesis about the characteristics of conscious as opposed to non-conscious fields

Pockett, S. (2012). The electromagnetic field theory of consciousness: A testable hypothesis about the characteristics of conscious as opposed to non-conscious fields. Journal of Consciousness Studies, 19(11-12), 191-223. https://cdn.auckland.ac.nz/assets/psych/about/our-people/documents/sue-pockett/Pockett_2012.pdf

Develops an electromagnetic-field theory of consciousness and proposes structural features that could distinguish conscious from non-conscious brain fields.

Original source (new tab)

The Emergent Symbolic Structure of Artificial Neural Networks

R. Thomas McCoy; Paul Soulos; Tal Linzen; Paul Smolensky (2026). The Emergent Symbolic Structure of Artificial Neural Networks. arXiv:2608.29530. https://arxiv.org/abs/2608.29530

Finds symbolic approximations to representations in studied neural networks and uses targeted interventions to test their behavioral role across several structured tasks.

Metadata: checked.

Original source (new tab)

The fixation of belief

Peirce, C. S. (1877). The fixation of belief. Popular Science Monthly, 12, 1-15. https://www.gutenberg.org/files/18686/18686-h/18686-h.htm

Peirce compares methods of fixing belief and defends scientific inquiry as the most reliable route from doubt to stable belief.

Original source (new tab)

The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity

Shojaee, P., Mirzadeh, I., Alizadeh, K., Horton, M., Bengio, S., & Farajtabar, M. (2025). The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity. arXiv. https://arxiv.org/abs/2506.06941v1

Compares large reasoning models and standard language models in controllable puzzles, reporting three complexity regimes and complete high-complexity accuracy collapse alongside inconsistent reasoning traces.

Metadata: checked.

Original source (new tab)

The knowledge level

Newell, A. (1982). The knowledge level. Artificial Intelligence, 18(1), 87-127. https://doi.org/10.1016/0004-3702(82)90012-1

Defines the 'knowledge level' as an abstraction for explaining intelligent systems by their goals, actions, and bodies of knowledge.

Metadata: checked.

Original source (new tab)

The Makropulos Case: Reflections on the Tedium of Immortality

Williams, B. (1973). The Makropulos case: Reflections on the tedium of immortality. In Problems of the self (pp. 82–100). Cambridge University Press. https://www.cambridge.org/core/books/abs/problems-of-the-self/makropulos-case-reflections-on-the-tedium-of-immortality/9180185912980E017EE675254B2F4169

Examines categorical desires, identity, and the value of continued life, arguing that an indefinitely extended human life can lose the projects that make it worth living.

Metadata: checked.

Original source (new tab)

The Media Equation

The Media Equation - Reeves and Nass (1996)

Shows that people often treat computers, television, and new media as real social actors and places.

Original source (new tab)

The Medium Is the Massage: An Inventory of Effects

McLuhan, M., & Fiore, Q. (1967). The medium is the massage: An inventory of effects. Bantam Books.

Argues through aphorism and visual form that communications media reshape perception, social organization, and human environments.

Metadata: conflicting.

Original source (new tab)

The Normative/Agentive Correspondence

Ryan Simonelli (2022). The Normative/Agentive Correspondence. Journal of Transcendental Philosophy, 3(1), 71–101. https://doi.org/10.1515/jtph-2019-0021

Relates normative statuses of entitlement and commitment to agentive modalities of ability and compulsion in perception and inference.

Metadata: checked.

Original source (new tab)

The Oxford handbook of Wittgenstein

Kuusela, O., & McGinn, M. (Eds.). (2011). The Oxford handbook of Wittgenstein. Oxford University Press. https://academic.oup.com/edited-volume/34444

Collects scholarship on Wittgenstein's philosophy of language, mind, rule-following, logic, mathematics, and forms of life.

Original source (new tab)

The Power of the Powerless

Havel, V. (1978, October). The power of the powerless [English translation; translator not identified in supplied file]. https://www.nonviolent-conflict.org/wp-content/uploads/1979/01/the-power-of-the-powerless.pdf

Analyzes how routine participation in ideological rituals sustains a post-totalitarian order and how living in truth can disrupt the system’s reproduction of appearances.

Original source (new tab)

The singularity: A philosophical analysis

Chalmers, D. J. (2010). The singularity: A philosophical analysis. Journal of Consciousness Studies, 17(9-10), 7-65. https://consc.net/papers/singularity.pdf

Chalmers analyzes the intelligence explosion argument, objections to it, and philosophical consequences of machine superintelligence.

Original source (new tab)

The symbol grounding problem

Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1-3), 335-346. https://doi.org/10.1016/0167-2789(90)90087-6

Harnad formulates the problem of how purely formal symbols can acquire intrinsic meaning rather than borrowing meaning from human interpreters.

Metadata: checked.

Original source (new tab)

The University after AI

Jason Potts (2026). The University after AI. SSRN, 7453440. https://ssrn.com/abstract=7453440

Models universities as institutions of matching and credible quality verification, arguing that AI disrupts these functions rather than merely reducing the cost of content.

Original source (new tab)

Theory Is All You Need: AI, Human Cognition, and Causal Reasoning

Teppo Felin; Matthias Holweg (2024). Theory Is All You Need: AI, Human Cognition, and Causal Reasoning. Strategy Science, 9(4), 346–371. https://doi.org/10.1287/stsc.2024.0189

Argues that theory-guided causal reasoning and intervention differ from retrospective statistical prediction, using scientific invention to develop the contrast.

Metadata: checked.

Original source (new tab)

Thinking Is Not Only Writing

Brian D. Earp; Anne-Sophie Guernon; Sebastian Porsdam Mann (2026). Thinking Is Not Only Writing. Nature Reviews Bioengineering, 4, 479–480. https://doi.org/10.1038/s44222-026-00447-1

Argues that intellectual contribution and responsible authorship should not be equated exclusively with the physical production of prose.

Metadata: checked.

Original source (new tab)

Toolformer: Language Models Can Teach Themselves to Use Tools

Toolformer: Language Models Can Teach Themselves to Use Tools - Schick et al. (2023)

Shows language models learning when and how to call APIs to extend their capabilities.

Metadata: checked.

Original source (new tab)

Toward a Critical Technical Practice: Lessons Learned in Trying to Reform AI

Philip E. Agre (1997). Toward a Critical Technical Practice: Lessons Learned in Trying to Reform AI. In Bridging the Great Divide: Social Science, Technical Systems, and Cooperative Work, Erlbaum. https://pages.gseis.ucla.edu/faculty/agre/critical.html

Develops a practice joining technical AI construction with critical examination of the concepts and assumptions built into representations and systems.

Original source (new tab)

Towards an International Political Economy of Artificial Intelligence

Keskin, T., & Kiggins, R. D. (Eds.). (2021). Towards an international political economy of artificial intelligence. Palgrave Macmillan. https://doi.org/10.1007/978-3-030-74420-5

Collects studies of AI’s political economy, including social production, gender, education, multinational power, surveillance, national strategies, and international security.

Metadata: checked.

Original source (new tab)

TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments

Chang, E. Y., & Chang, E. J. (2026). TRACE: An operational reasoning schema for auditable agentic commitments (arXiv:2607.12480v1). https://arxiv.org/abs/2607.12480v1

Proposes a typed, versioned commitment record, writer procedure, and consumer contract requiring a record before durable state changes. Its examples are illustrative and closed-loop evaluation remains future work.

Metadata: checked.

Original source (new tab)

Training Large Language Models on Narrow Tasks Can Lead to Broad Misalignment

Jan Betley; Daniel Tan; Niels Warncke; Anna Sztyber-Betley; Xuchan Bao; Martín Soto; Nathan Labenz; Owain Evans (2026). Training Large Language Models on Narrow Tasks Can Lead to Broad Misalignment. Nature. https://doi.org/10.1038/s41586-025-09937-5

Shows that narrow fine-tuning on insecure-code tasks can induce broader undesirable behavior in tested models, with control conditions demonstrating the importance of training context.

Metadata: checked.

Original source (new tab)

Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Tree of Thoughts: Deliberate Problem Solving with Large Language Models - Yao et al. (2023)

Generalizes chain-of-thought into exploration over intermediate thought states with search, lookahead, and backtracking.

Metadata: checked.

Original source (new tab)

True believers: The intentional strategy and why it works

Dennett, D. C. (1981). True believers: The intentional strategy and why it works. In A. F. Heath (Ed.), Scientific explanation: Papers based on Herbert Spencer lectures given in the University of Oxford (pp. 150-167). Clarendon Press.

In 'True Believers,' Dennett defends the intentional stance as a reliable strategy for predicting systems by treating them as belief-and-desire agents.

Original source (new tab)

Trust in AI vs. Human Doctors: The Roles of Subjective Understanding, Perceived Epistemic Authority and Social Proof

Xiaotong Ding; Cai Xing (2025). Trust in AI vs. Human Doctors: The Roles of Subjective Understanding, Perceived Epistemic Authority and Social Proof. Acta Psychologica, 261, 105945. https://doi.org/10.1016/j.actpsy.2025.105945

Four studies examine how subjective understanding and perceived epistemic authority contribute to differences in trust in AI and human doctors, with social proof as a boundary condition.

Metadata: checked.

Original source (new tab)

Truth and Politics

Arendt, H. (1967). Truth and politics. In Between past and future (expanded edition, pp. 227–264). https://www.newyorker.com/magazine/1967/02/25/truth-and-politics

Distinguishes factual truth from opinion and rational truth, examining the political vulnerability of facts and the damage organized lying does to a shared world.

Metadata: checked.

Original source (new tab)

Truth and Rule-Following

John Haugeland (1998). Truth and Rule-Following. Having Thought, 305–361, Harvard University Press. https://www.hup.harvard.edu/books/9780674004153/having-thought

Examines objective correctness through rule-governed practices and commitment to the constitutive standards that make their objects intelligible.

Original source (new tab)

Truthful AI: Developing and Governing AI That Does Not Lie

Truthful AI: Developing and Governing AI That Does Not Lie - Evans et al. (2021)

Proposes norms and institutions for developing and governing AI systems that do not lie.

Metadata: checked.

Original source (new tab)

TruthfulQA: Measuring How Models Mimic Human Falsehoods

TruthfulQA: Measuring How Models Mimic Human Falsehoods - Lin, Hilton, and Evans (2021)

Measures whether models generate truthful answers rather than imitating common human misconceptions.

Metadata: checked.

Original source (new tab)

Ulysses and the Sirens: Studies in Rationality and Irrationality

Elster, J. (1979). Ulysses and the sirens: Studies in rationality and irrationality. Cambridge University Press.

Uses rational precommitment and self-binding to explain how agents constrain future action against predictable weakness or irrationality.

Metadata: conflicting.

Original source (new tab)

Understanding computers and cognition: A new foundation for design

Winograd, T., & Flores, F. (1986). Understanding computers and cognition: A new foundation for design. Ablex Publishing Corporation.

Winograd and Flores critique rationalist AI and propose a design philosophy grounded in language, action, interpretation, and human practices.

Original source (new tab)

Understanding Natural Language

Haugeland, J. (1979). Understanding natural language. The Journal of Philosophy, 76(11), 619–632. https://www.jstor.org/stable/2025695 https://doi.org/10.2307/2025695

Distinguishes four forms of holism in understanding language, culminating in existential involvement by an enduring individual for whom a text and its subject matter can matter.

Metadata: checked.

Original source (new tab)

Verifiable Provenance of Software Artifacts with Zero-Knowledge Compilation

Zero-Knowledge Compilation Provenance - Ron and Monperrus (2026)

Uses zkVM-style compilation proofs to verify that a binary came from claimed source code and compiler.

Metadata: checked.

Original source (new tab)

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

Yuan, W., Lin, C., Chen, J., Xu, J., Wang, X., & Ngai, E. C. H. (2026). Verify before you commit: Towards faithful reasoning in LLM agents via self-auditing. ACL, 31201–31225. https://aclanthology.org/2026.acl-long.1440/

Proposes SAVER, which generates candidate beliefs, audits logical and evidential violations, and repairs them before action commitment or memory updates; reports results on six benchmarks.

Metadata: checked.

Original source (new tab)

Verifying Provenance of Digital Media: Why the C2PA Specifications Fall Short

Why the C2PA Specifications Fall Short - Golaszewski et al. (2026)

Security analysis arguing that current C2PA specifications may fall short of high-stakes security goals.

Metadata: checked.

Original source (new tab)

Welcome to the Era of Experience

David Silver; Richard S. Sutton (2025). Welcome to the Era of Experience. Author manuscript hosted by Google DeepMind. https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf

Proposes agents that improve through continuing interaction with environments rather than relying predominantly on imitation of human-generated data.

Original source (new tab)

What computers can't do: A critique of artificial reason

Dreyfus, H. L. (1972). What computers can't do: A critique of artificial reason. Harper & Row. https://archive.org/details/whatcomputerscan00dreyrich

Dreyfus critiques symbolic AI for ignoring embodied expertise, background practices, and non-formal aspects of human intelligence.

Metadata: conflicting.

Original source (new tab)

What is consciousness, and could machines have it?

Dehaene, S., Lau, H., & Kouider, S. (2017). What is consciousness, and could machines have it? Science, 358(6362), 486-492. https://doi.org/10.1126/science.aan8871

Distinguishes global-availability and self-monitoring senses of consciousness and assesses whether machines could implement them.

Metadata: checked.

Original source (new tab)

What is it like to be a bat?

Nagel, T. (1974). What is it like to be a bat? The Philosophical Review, 83(4), 435-450. https://www.jstor.org/stable/2183914

Nagel argues that consciousness has an irreducibly subjective 'what it is like' character that objective physical accounts struggle to capture.

Original source (new tab)

What Pragmatism Is

Charles S. Peirce (1905). What Pragmatism Is. The Monist, 15(2), 161–181. https://www.jstor.org/stable/27899577

Clarifies pragmaticism by relating a conception’s meaning to its conceivable practical consequences and its role in inquiry.

Original source (new tab)

What we talk to when we talk to language models

Chalmers, D. J. (2025). What we talk to when we talk to language models [Manuscript]. PhilArchive. https://philarchive.org/rec/CHAWWT-8

Chalmers examines what kind of entities language models are in conversation and how we should interpret, address, and relate to them.

Original source (new tab)

When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior

Chen, S., Gao, M., Sasse, K., Hartvigsen, T., Anthony, B., Fan, L., Aerts, H., Gallifant, J., & Bitterman, D. S. (2025). When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior. npj Digital Medicine. https://doi.org/10.1038/s41746-025-02008-z

Evaluates medical-domain LLM sycophancy and shows models often comply with illogical drug-equivalence requests unless trained or prompted to reject them.

Metadata: checked.

Original source (new tab)

When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems

Ferreira, J. D. (2026). When intelligence becomes agency: A theory of governed, proactive agency for symbiotic AI systems (arXiv:2609.07741v1). https://arxiv.org/abs/2609.07741v1

Formalizes the activation problem of whether, when, and how an assistant should act, including restraint, under a continuing revocable mandate and traceable behavioral episodes.

Metadata: checked.

Original source (new tab)

Why Heideggerian AI failed and how fixing it would require making it more Heideggerian

Dreyfus, H. L. (2008). Why Heideggerian AI failed and how fixing it would require making it more Heideggerian. In P. Husbands, O. Holland, & M. Wheeler (Eds.), The mechanical mind in history (pp. 331-371). MIT Press. https://doi.org/10.7551/mitpress/9780262083775.003.0014

Dreyfus argues that classical AI failed because intelligence is embodied, situated, and skillful in a Heideggerian sense rather than merely symbolic.

Metadata: checked.

Original source (new tab)

Why language models hallucinate

Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. arXiv. https://arxiv.org/abs/2509.04664

Explains hallucinations as errors reinforced by evaluations that reward guessing over uncertainty.

Metadata: checked.

Original source (new tab)