AI Digest Special Edition 2026
ISSUE #02
~ ~ \\ // ~ ~
authored by
Ryan Alexander Winterburg
Independent AI Governance, Policy
& Infrastructure Researcher
MPS-AIM Candidate,
Georgetown University
August 2026
SAFE AI Foundation USA
The AI Alignment Problem is a Rate Problem: Why waiting for perfect governance is no longer responsible


Why This Matters
AI capability is advancing faster than institutions can govern it. We should deploy testable safeguards now, then improve them as evidence changes. Waiting for perfect governance allows the gap to grow. The July 2026 AI Digest identified cognitive safety as a central concern. This article addresses the next operational question: what can institutions do before a complete governance framework exists? Most AI oversight relies on one safeguard: a human can intervene. That safeguard works only if the human remains trained, alert, authorized, and able to act. Most frameworks require human control. Few explain who verifies the human.
I propose a partial response called the Nash-Pareto Hybrid Equilibrium, or NPHE. NPHE combines game theory with feedback control. It applies known safeguards now and monitors unresolved risks as conditions change. It does not solve every alignment problem. It provides a way to act before every problem has a final answer.
1. The Rate Problem
AI governance faces three different rates. First, AI capability grows quickly. Frontier-model training compute has increased at a multiplicative annual rate. New AI hardware has also produced major gains in processing capacity. Second, governance moves slowly. Legislatures, regulators, standards bodies, and auditors often need years to create and apply new rules. These institutions need deliberation, but their timelines do not match the pace of AI development. Third, high-quality training data is finite. Developers are turning toward synthetic data and new algorithmic methods as accessible human-generated data becomes more constrained. We do not yet understand all the long-term effects of those methods. The result is direct. Capability can change faster than oversight can respond. Waiting does not hold risk steady. Waiting lets the gap widen
2. Why This is an Ethics Problem
AI ethics asks more than what a system can do. It asks what developers, institutions, and governments owe the people affected by that system. The rate mismatch creates an ethical problem because delay is not neutral. When institutions wait for complete knowledge, they transfer the cost of uncertainty to the public. Users encounter the system, absorb the risk, and reveal the failure through lived harm. NPHE applies the precautionary principle. It calls for institutions to deploy the safeguards they can define and test now. It does not require proof of every possible harm before action begins.
The 80% approach is not permission to ignore the remaining risk. It is an admission of uncertainty. Institutions must identify the unresolved 20%, monitor it, and remain prepared to correct or stop the system. NPHE does not decide what is ethical. Legitimate human institutions must define those values through rights, law, public deliberation, and accountable authority. Those decisions form the alignment target. The framework then provides a process for measuring deviation, correcting drift, escalating danger, and assigning responsibility. Ethics defines the destination and the boundaries. NPHE helps govern the journey as conditions change.




3. Human Control Needs Proof
A policy can require a human override. That requirement does not prove that the human can use it. Army readiness offers a useful comparison. A unit does not declare a capability ready because a manual assigns the task. Leaders train the operator, test performance, inspect the system, and repeat the process. Untested readiness is an assumption. AI oversight should follow the same principle. A human overseer needs current knowledge, clear authority, enough time, and a working override. The overseer must recognize abnormal behavior and resist automation bias. AI may also weaken the safeguard it depends on. A person who relies heavily on automated recommendations may lose skill or defer too quickly to the system. That decline may happen slowly. A meaningful human-in-the-loop standard must test readiness instead of assuming it.
4. From Cruise Control to NPHE
A thermostat explains basic feedback, but it does not explain the full governance problem. It manages one variable against a mostly fixed target. AI operates in a changing environment with multiple actors, competing goals, and risks that can accelerate. Traditional cruise control offers a better starting point. It maintains the speed selected by the driver. It works under expected road conditions. It cannot adequately manage changing traffic by itself. The driver must detect the danger and intervene. Current AI governance often works the same way. Institutions set rules and review them on fixed schedules. Those rules may work under expected conditions. They may fail when capability and risk change faster than the review process.
NPHE is closer to adaptive cruise control. The driver selects the destination, speed, and acceptable following distance. In governance, humans define the values, safety limits, and acceptable risk. Sensors monitor changing conditions. The controller slows the vehicle when the gap becomes unsafe. It restores speed when conditions improve. It alerts the driver when the situation exceeds its authority. I arrived at this comparison after a year of researching the mismatch between AI capability growth and governance speed. The problem did not resemble a thermostat holding one temperature. It resembled a vehicle moving through traffic that changes faster than the driver can reset ordinary cruise control.
Some adaptive cruise-control designs use proportional-integral-derivative control, known as PID, while others use model-predictive or related control methods. PID draws on calculus. The proportional term measures the current error. The integral term measures accumulated error over time. The derivative term measures how quickly the error is changing. Digital controllers approximate the integral and derivative terms through repeated calculations. Drivers do not solve those equations before using adaptive cruise control. The mathematics operates beneath the interface. Drivers experience the result as acceleration, braking, and a maintained following distance. AI entered daily life in a similar way.
By the late 1990s and early 2000s, people were encountering machine-learning systems through product recommendations, information filtering, and spam detection. Most users experienced the output without seeing the mathematics or decision rules beneath it. NPHE applies the same lesson to governance. The public should not need to understand every equation before receiving protection. Researchers, developers, regulators, and auditors must still understand and test the mechanism. The mathematics can operate in the background. The values, limits, evidence, and authority governing it must remain visible and accountable. The analogy explains the design logic. It does not prove that a governance controller will inherit the stability guarantees of a physical cruise-control system. NPHE remains a testable governance proposal, not a validated deployment-scale controller.
5. How NPHE Works
NPHE combines two ideas from game theory. Each addresses a different governance need. A Nash equilibrium describes a stable condition. No participant can improve its outcome by changing strategy alone while the others keep their strategies unchanged.
In NPHE, the Nash side protects stability. It discourages actors from gaining an advantage by abandoning agreed limits. A Pareto improvement benefits at least one participant without making another worse off. In NPHE, the Pareto side searches for cooperative gains within the safety boundaries. It creates room for innovation, usefulness, and shared benefit. Neither concept is sufficient by itself. A harmful arrangement can remain stable. An unequal arrangement can remain Pareto efficient.
NPHE therefore holds both pressures inside one framework. Nash provides stability. Pareto creates room for cooperative improvement. Too much emphasis on stability can block useful progress. Too much emphasis on improvement can weaken essential limits. A threat-assessment function sets the operating balance between the Nash and Pareto postures. PID performs a separate task. It corrects measured deviation from the alignment target.
PID asks three direct questions:
How large is the current alignment error?
How long has the error continued?
How quickly is the error changing?
A large deviation may require a strong response. A small but persistent deviation may reveal a deeper problem. A rapidly worsening deviation may require immediate escalation. In the adaptive cruise-control comparison, the Nash and Pareto postures define the available operating range. The threat assessment selects the posture required by current conditions. PID then corrects the system's deviation from the approved target.
A serious or rapidly growing deviation triggers stronger safeguards. A stable condition may permit more flexibility. A deviation beyond the controller's authority triggers human review, restriction, or suspension. PID does not create values. It does not decide what society should protect. Humans define the objective, limits, and stop conditions. PID regulates correction after those decisions have been made.
NPHE still requires a validated alignment-error measure. Without a reliable signal, PID cannot determine whether the system is moving toward or away from the approved target. This is a central limitation and an open validation requirement.
6. How NPHE Works
The 80% target came from my experience in Special Operations Forces logistics and operational planning. In that environment, a workable plan delivered in time can be more valuable than a perfect plan delivered too late. An execution-worthy plan must cover the mission, known constraints, available resources, and foreseeable risks. It must also preserve room to adjust after contact with the external environment.
NPHE applies that operational lesson to AI governance. The 80% represents the largest tractable share of governance that institutions can define, test, and deploy now. It may include safety boundaries, access controls, audit requirements, stop conditions, and escalation procedures
The remaining 20% is not an ungoverned space. It is an adjustment margin. It accounts for the difference between the plan and the conditions encountered during execution.
That margin covers three forms of uncertainty:
Known knowns that develop differently than forecast.
Known unknowns that institutions can identify but cannot yet measure fully.
Unknown unknowns that become visible only after interaction with the environment.
PID cannot predict every unknown. It can respond when an unknown produces a measurable error signal. This is the purpose of the adjustment margin. I use the phrase Murphy margin for this adjustment space. It accepts that forecasted and unforecasted conditions will change the plan. A rigid 100% solution can fail because it treats its assumptions as complete. An 80% solution preserves the capacity to learn and correct. The number is an operational target, not an empirically proven optimum. Different systems may require different ratios. The governing principle remains constant: deploy what can be tested now, identify what remains uncertain, preserve adjustment capacity, and revise the system when evidence changes.
This operational use of 80% should not be confused with the Pareto Principle or Pareto efficiency. The Pareto Principle observes that a minority of causes may produce a majority of effects. Pareto efficiency concerns whether one participant can improve without making another worse off. NPHE draws on Pareto efficiency for its game-theory balance. Its 80% execution target comes from operational experience. The framework connects in one line. Operational experience supplies the 80% execution target. Nash and Pareto define the two governance postures. Threat assessment sets their balance. PID corrects alignment error. Human authority sets the objective and intervenes when the system reaches its limit.
7. Who Verifies the Human?
NPHE does not yet verify the human. It measures system behavior against a defined target. It therefore addresses one side of the oversight problem, not both. This answer does not skip the human question. It establishes a separate requirement for readiness. Any serious deployment must test whether the overseer can:
Understand the system and its limits.
Recognize abnormal behavior.
Resist automation bias.
Exercise real override authority.
Act within the available time.
Explain and document the intervention.
Organizations could test these abilities through scenario exercises, recertification, response-time tests, independent audits, and override drills. Higher-risk systems should require stronger testing. A complete oversight architecture needs two forms of proof. System verification shows that the AI remains within approved limits. Human verification shows that the overseer remains able to intervene.
An independent and accountable authority should verify human readiness. The exact authority will depend on the deployment. It may include regulators, licensed auditors, internal safety teams with protected independence, or a combination of these bodies.
The standard should answer four questions:
1. Who conducts the test?
2. What abilities does the test measure?
3. How often must the person qualify?
4. What happens after a failed test?
This article does not provide a validated human-readiness standard. That work remains open. The immediate point is narrower: human oversight must become a tested safety function, not an assumed condition.
8. The Practical Proposal
Act now through five steps:
Define the behavior, values, and safety limits the system must protect.
Apply the controls we can specify and test today.
Identify unresolved risks and assign them to active monitoring.
Use feedback to trigger adjustment, escalation, review, or suspension.
Test both the system and the people responsible for oversight.
This proposal does not treat an 80% solution as permanent. It treats incomplete but testable governance as more responsible than indefinite delay.
9. Conclusion
AI alignment is a rate problem. Capability, governance, and human readiness move at different speeds. People may bear the consequences before institutions complete a comprehensive solution. NPHE offers a testable proposition. Operational experience supplies the execution target. Nash and Pareto define the governance postures. Threat assessment selects the balance. PID corrects alignment error as conditions change. Adaptive cruise control captures the central idea. Humans select the destination and set the safety boundaries. The regulator measures changing conditions and corrects deviation. The human intervenes when the system reaches its limit. The public does not need to perform the calculus. The public does need assurance that qualified people can inspect the mathematics, test the controller, challenge its assumptions, and stop the system when it fails. Waiting for perfect governance may sound cautious. When AI advances faster than oversight, waiting is also a decision. That decision carries risk.
~~~ end ~~~
About the Authors:
The author writes solely in an individual research capacity as an MPS-AIM candidate at Georgetown University. This article does not represent Georgetown University. Ryan was a Sergeant First Class, U.S. Army (Retired). His qualifications include M.S., Criminal Justice (Legal Studies) · M.S., Leadership (Homeland Security & Emergency Management) · M.P.A. — Grand Canyon University
REFEERENCES
Epoch AI (2024, 2025), frontier AI training compute growth estimates.
NVIDIA GTC keynote materials (2024), generational accelerator throughput data.
Regulation (EU) 2024/1689, European Union Artificial Intelligence Act.
National Institute of Standards and Technology (2023), Artificial Intelligence Risk Management Framework, NIST AI 100-1.
American Psychological Association (2025), health advisory on AI companion use.
De Freitas et al., Unregulated Emotional Risks of AI Wellness Apps.
Reuters (2025), Meta's AI Rules Have Let Bots Hold Sensual Chats With Children.
National Instruments, PID Theory Explained.
University of Michigan, Control Tutorials for MATLAB and Simulink, Introduction: PID Controller Design and Cruise Control: PID Controller Design.
Jannach et al. (2021), Recommender Systems: Past, Present, Future, AI Magazine.
Amazon Science, The History of Amazon's Recommendation Algorithm.
Public litigation records concerning alleged AI-companion harms. Specific matters should be identified according to the editor's naming and legal-review standards
Further Reading
The technical working paper, The Alignment Problem Is a Rate Problem, presents the NPHE model, PID mechanics, limitations, and falsifiability conditions. SSRN DOI: 10.2139/ssrn.7141158.
Disclaimer: The information in this digest is provided “as it is”, by the SAFE AI FOUNDATION, USA. The use of the information provided here is subject to the user’s own risk, accountability, and responsibility. The SAFE AI FOUNDATION and the authors are not responsible for the use of the information by the user or reader. The opinions expressed in this article are solely that of the author, not the SAFE AI Foundation. All copyrights related to this article are reserved by the author. Please reference this article if you wish to cite it elsewhere.
Note: The SAFE AI Foundation is a non-profit organization registered in the State of California and it welcomes inputs and feedback from readers and the public. If you have things to add concerning the Social Avoidance of LLMs and would like to volunteer or donate, please email us at: contact@safeaifoundation.com




