Belfer Center for Science and International Affairs
Permanent URI for this communityhttps://dash.harvard.edu/handle/1/42731987
Browse
Search Results
Publication Architecting Trust: A Modular Framework for the Operational Deployment of Autonomous Systems
(2026) Pederson, JoelFor the United States and its allies, the incorporation of non-deterministic Artificial Intelligence (AI) and Machine Learning (ML) systems into tactical platforms presents significant challenges. Certifying these agents for safety-critical operations remains a considerable barrier to their deployment and widespread adoption. This paper examines how AI-enhanced tactical solutions can be effectively fielded despite the inherent risks involved. The analysis begins by discussing the history of automation bias in weapons systems and the novel vulnerabilities introduced by AI/ML-powered solutions. The challenge with these emerging techniques is that unlike traditional software, these algorithms present risks throughout their lifecycle. From adversarial perturbations that can be introduced during model training to stochastic failures during operation, the risks associated with AI/ML-based algorithms cannot simply be mitigated by the legacy safety systems embedded in platforms today. While traditional “physics-based” guardrails (e.g., automated ground collision avoidance) effectively prevent kinematic disasters, they on their own lack the sophistication to address the cognitive and perception errors inherent to modern AI. To safely proliferate AI/ML solutions, a new safety-focused reference architecture is required. This paper proposes using a modular “Safety Sidecar” architecture that operates across the software’s lifecycle. The research defines a strategy that acquisition authorities can consider building from to systematically embed safety barriers directly into the system’s training, perception, and planning loops. The framework outlines how the training of models can be protected and refined for operational fielding through a structured lifecycle assurance approach. Specifically, the architecture introduces a “Perception Gatekeeper” which validates input integrity against adversarial or degraded sensor data in real-time. The framework also integrates algorithms running operationally in the loop with AI/ML models to function as a “Model Constraint Guardian ” to enforce safety barriers directly into the system’s perception and planning loops. The Safety Sidecar concept can function as a unified framework for AI/ML lifecycle assurance. The approach can facilitate effective integration of safety into both developmental and operational phases of system development. This contribution provides a structured pathway to help ensure that safety is a continuous property from model training through to battlefield deployment.
Publication De-Risking Defense Innovation at the Earliest Stages: The Strategic Role of Entrepreneurial Fellowships and Early Non-Dilutive Grant Funding
(2026) Kennedy, Elizabeth; Emmi, LaurenDespite billions in sustained U.S. investment in defense-related basic and applied research, many high-potential technologies fail to transition from the laboratory to operational capability, resulting in capital inefficiency, delayed warfighter access, and erosion of U.S. technological advantage. This gap is most pronounced at the earliest technology readiness levels (TRLs), where technical uncertainty, unclear demand signals, and limited access to capital stall progress. Traditional defense acquisition and grant mechanisms are optimized for integrating mature technologies against well-defined requirements, rather than high-risk early technologies, leaving early-stage innovations without viable pathways to maturation and adoption. This paper examines entrepreneurial fellowship models, defined as structured, multi-year programs that provide early-stage, non-dilutive funding, as critical infrastructure for defense and other critical innovation. These models provide intensive mentorship through company formation, salary support enabling full-time founder commitment, and early connections to investors and government customers. By sustaining founders with modest grant funding through the highest-risk period of financing and technology validation—before conventional acquisition or venture capital funding are viable—entrepreneurial fellowships prevent loss of critically important technological innovations before they reach acquisition relevance. Beyond addressing this core failure in the innovation pipeline, they enable earlier alignment with military requirements while achieving public capital efficiency. Drawing on evidence from defense-adjacent fellowship programs and comparable innovation initiatives, this paper argues that early-stage grant-based fellowships outperform other funding mechanisms in three key ways: (1) accelerating transition from proof-of-concept to defensible prototype, (2) reducing downstream acquisition risk by enabling earlier validation of operational use cases and requirements, and (3) crowding in private capital at later stages without prematurely forcing commercial scaling or misaligned market entry. Importantly, these programs complement—not replace—existing acquisition pathways by feeding them with better-defined, more advanced technologies. These effects are particularly pronounced in strategically critical technology domains, including position, navigation, timing (PNT) capabilities, advanced manufacturing, biomanufacturing, and advanced materials, where fellowship-supported companies have delivered novel technologies that surpass existing solutions in both capability and cost. In an era in which future conflicts will be decided by technological superiority, the nation that validates and fields critical technologies first accrues decisive strategic advantage. This paper concludes with policy recommendations for the Department of Defense (DoD) and interagency partners to expand and institutionalize early-stage entrepreneurial fellowship funding as a strategic tool for strengthening the defense industrial base, shortening time-to-capability, and preserving U.S. technological advantage amid intensifying global competition.
Publication Escalation Protocols for Autonomous Naval Vessels: A Legal-Technical Framework for Safe and Credible Maritime Autonomy
(2026) Foley, JordanAutonomous and remotely supervised surface vessels are moving from pilots to operational deployment across defense and commercial shipping. U.S. Navy unmanned surface vessels (USVs)—including mine countermeasures platforms—and commercial Maritime Autonomous Surface Ships (MASS) operate in congested waterways alongside human mariners yet remain governed by legal regimes built around human signaling and post-incident accountability.12 This paper argues that maritime autonomy faces a governance vulnerability: systems may navigate safely but remain legally brittle if their conduct cannot be rendered intelligible under the standards used to assign responsibility after incidents. Maritime compliance is evaluated through reconstruction. The COLREGs apply to “all vessels,” require standardized lights and sound signals, and permit supplemental warnings under Rule 36 only if they cannot be mistaken for authorized signals.3 UNCLOS reinforces compliance as sovereign responsibility by requiring flag states to ensure conformity with generally accepted international safety regulations.4 These rules create an evidentiary burden autonomy often cannot meet. Collision doctrine underscores the stakes: statutory navigation violations can trigger presumptions of fault that must be rebutted by proof.5 In coercive or ambiguous encounters, the burden intensifies because escalation governance depends on demonstrable sequencing and restraint rather than after-action narrative.6 Existing maritime systems generate abundant data but rarely preserve a unified, time-synchronized record of what an autonomous vessel perceived, what it signaled, how it maneuvered, and whether human-on-the-loop judgment boundaries were maintained.7 The paper proposes a maritime assurance operating layer as the enabling infrastructure for lawful autonomy at scale. Non-controlling and platform-agnostic, it does not steer the vessel or execute coercive actions; it fuses heterogeneous sensor streams and operational logs into an audit-grade timeline that captures signaling and escalation sequencing, including Rule 36 warnings and human intervention pathways. By making compliance reconstructable and restraint provable, assurance infrastructure reduces legal exposure, narrows narrative ambiguity in contested waters, and supports coalition and commercial adoption through portable accountability outputs. The core conclusion is that maritime autonomy will scale only if it becomes legible: evidence-grade assurance is a condition of credible, stabilizing autonomy at sea.
Publication When Bans Don't Build Markets: Rethinking the DoD’s Critical Mineral Strategy
(2026) Russell, BethanyThe U.S. government and the Department of Defense (DoD) are increasingly treating critical minerals as a core national security priority. Over the past several years, both entities have invested heavily in supply-side interventions—financing .projects, supporting processing capacity, and expanding mineral stockpiles. Far less attention however has been given to the design and sequencing of demand-side policies intended to create a durable market for U.S. and allied-sourced minerals. One of the DoD’s most consequential demand-side measures is its 2027 ban on procuring certain covered minerals and components sourced from China. The intent of the ban is clear: reduce dependencies on the People’s Republic of China (PRC) and catalyze a secure supply chain. However, as it is currently structured, the ban risks outpacing industrial realities. It imposes a compliance deadline before sufficient alternative mining, refining, and processing capabilities likely exist. It targets minerals that have received comparatively limited DoD investment. It may raise downstream acquisition costs without corresponding budget adjustments. It disproportionately burdens smaller defense firms, potentially narrowing competition within the defense industrial base. The central argument of this essay is not that demand-side intervention is misguided. Demand-side policies are crucial for the success of the DoD’s supply-side investments. Rather, the current approach is poorly synchronized. A ban alone does not create a market; it creates a compliance obligation. Without aligned capital investment, pricing mechanisms, and scale beyond DoD procurement, the ban risks producing bottlenecks, volatility, and consolidation rather than resilience.
Publication Frozen, Blind, and Air-Gapped: Reconciling Frontier Model Capabilities with Defense Infrastructure Realities
(2026) Varshavsky, Steven; Krishnan, NaveenThe DoD's current posture is one of importation - forcing commercial models into defense infrastructure. This works in permissive environments and fails structurally at the classified and tactical edge. A fine-tuned SLM on a low-SWaP IL6- certified device delivers more combat power than a frontier model in a CONUS cloud enclave, regardless of benchmark scores. Sustainable Al advantage comes from reproducible infrastructure. Maven, GIDE, and Agile Flag all succeeded as pilots but never scaled because the certification pathways, hardware pipelines, and data architecture to replicate them didn1 exist. Allied nations - NATO DIANA, UK MoD, Dstl - independently arrive at the same answer: small, certifiable, edge-viable models. The mismatch is infrastructural, not algorithmic. The tools and authorities exist. The task is alignment.
Publication Detecting Systematic Infrastructure Attacks via Geospatial Intelligence
(2026) Chen, Kevin; Griffiths, ColeModern warfare increasingly targets infrastructure indispensable to civilian survival, particularly water and food systems, as a tactical instrument of coercion. Despite extensive reporting, a capability gap persists in scalable tools for real-time damage detection and humanitarian impact prediction. This paper introduces an integrated geospatial artificial intelligence (GeoAI) framework leveraging multi-modal satellite imagery, computer vision, and data fusion to map vulnerabilities in infrastructure and supply chains. By fusing pixel-level change detection with ACLED conflict data and WFP logistics telemetry, we generate spatially explicit predictive risk indicators. We evaluate this toolkit through case studies on the destruction of the water-energy nexus in Ukraine and the food system blockades in Yemen. Contributions include a scalable OSINT toolkit, a multi-tiered early-warning typology, and a prototype dashboard architecture for crisis management. Finally, we address ethical considerations regarding demographically identifiable information (DII), demonstrating how bridging GeoAI with humanitarian analysis advances early-warning capabilities for national security and global response.
Publication The End of the Gray Zone? How AI-Enabled Cyber Rivals Kinetic Capabilities
(2026) Bahrami, Daria; Chang, Amy; Kouremetis, Michael; Devendorf, Erich; Saade, TiffanyHistorically, cyber operations have failed to deliver the strategic lethality predicted by early theorists. While disruptive, case studies ranging from the Russian power grid attacks to Stuxnet demonstrate that digital effects are often temporary, reversible, and notoriously difficult to synchronize with cross-domain operations. This paper argues that Artificial Intelligence (AI) is fundamentally altering this calculus, collapsing the temporal asymmetry between "gray zone" competition and high-intensity conflict. By automating the speed, scale, and integration of offensive operations, AI provides the "kinetic-equivalent" impact necessary to compel state behavior and alter deterrence dynamics. This analysis employs a conceptual and case-based methodology to explore the trajectory of AI-enabled warfare. First, it identifies the "digital ceiling" of pre-AI cyber operations, analyzing why manual coordination across DIME (Diplomatic, Information, Military, Economic) instruments limited strategic utility. Second, it investigates how agentic AI and Large Language Models (LLMs) bridge the cyber-kinetic gap. Through an examination of autonomous vulnerability exploitation, the research highlights how AI-enabled campaigns can generate continuous, compounding destruction and economic attrition that rivals kinetic bombardment. By collapsing the time between attack and recovery, AI-driven campaigns can outpace an adversary's financial and technical restoration capacity, effectively rendering "reversible" cyber effects strategically permanent. Finally, the paper addresses the critical policy void created by these advancements. Current risk frameworks are static and ill-equipped to measure dynamic AI behaviors, while the democratization of open-weight models renders traditional supply-chain interdiction insufficient. The research concludes that the blurring of the civilian-combatant divide requires a radical shift in U.S. policy: moving from risk aversion to mandatory "resilience by design." This includes establishing minimum security baselines for the private sector and accepting that the democratization of AI capabilities requires a more aggressive posture in offensive testing. Ultimately, this paper posits that to maintain global negotiating power, national security strategy must treat AI-driven cyber conflict not as a support function, but as a primary determinant of geopolitical influence.
Publication Why Workforce Governance Is the Limiting Factor in National Security Innovation
(2026) Lorell, DesireeArtificial intelligence (AI) is rapidly being integrated into national security organizations, promising speed, efficiency, and enhanced decision-making. Yet AI adoption often prioritizes tools and technical capability over workforce and governance systems responsible for interpreting and validating and acting on AI-enabled outputs. Drawing on patterns identified through doctoral research on information and communication technology selection implementation in defense learning and operational environments, and insights from a national security AI fellowship. This paper examines how AI adoption reshapes epistemic authority inside organizations. Unlike prior “game-changing” technologies, AI produces answer-like outputs compressing deliberation timelines and increasing pressure on judgment. Without explicit training, doctrine, and workforce safeguards, this dynamic can erode accountability, distort readiness, and incentivize premature adoption. The analysis identifies three failure modes. First, workforce systems often lack the data fidelity and role clarity to determine who is qualified to rely on or contest AI outputs. Second, governance mechanisms are often introduced after acquisition, limiting their influence on system design and workforce preparation. Third, productivity narratives emphasize efficiency gains while underrepresenting the coordination and verification burden that emerge in operational contexts. This study advances a workforce-centered governance model for AI integration in national security. It emphasizes tiered AI exposure based on readiness, doctrinal separation between automation, decision support, and authority, and institutionalized “check-the-checker” norms that preserve human judgment and command accountability. By reframing AI adoption as a workforce and governance challenge, this study offers a practical path for aligning innovation with mission assurance in contested and high-stakes environments.
Technology & National Security JournalCollection