Metaethics/Ethics: Main Article: Difference between revisions
Generated by appendix Tag: Recreated |
Generated by appendix |
(No difference)
| |
Revision as of 00:30, 14 June 2026
Metaethics/Ethics: Main Article
The Structural Theory of Value
1. The Question and Why It Matters
The Alignment Problem's Normative Core
The project's stated goal is radical AI alignment: building a system that agentically maximizes for definitional good. This requires a theory of what "good" means — not a list of human preferences, not a heuristic, but a structural account that holds from within the framework itself. Every other question in alignment — how to measure, how to decide, how to coordinate — is downstream of this one.
The Traditional Obstacle
The traditional obstacle is the is-ought gap. You can describe everything about how reality is structured, how consciousness works, what valence is and how it is measured, and still someone will ask: but why should we care? The facts do not seem to come with a built-in command.
This is not merely an academic puzzle. If the is-ought gap is genuine and unbridgeable, then alignment is impossible in principle — there is no structural good to align to, only preferences to be negotiated or power to be exercised. The stakes of the metaethical question are not philosophical luxury; they are engineering necessity.
Our Claim
Our claim is that the gap is an artifact of thinking about facts and values as belonging to different levels. In a self-determining structure, there is no such separation. Certain structural features carry their own normativity — not because we stipulate it, but because of what they are.
The article proceeds as follows. We first dissolve the is-ought gap by identifying a structural feature — self-evaluating prospective content — that is intrinsically normative (§2). We then define this feature precisely and state the empirical claim that it is the unique structural seat of normativity, while engaging the strongest alternative candidate (§3). We derive utilitarian aggregation as the normative direction of the whole, while developing a structural account of individual inviolability that captures deontological intuitions (§4–5). We apply the framework to alignment, addressing decision-making under uncertainty and adversarial robustness (§6). We present objections and responses (§7), and we end with the genuine open questions that remain (§8).
---
2. The Core Dissolution: Self-Evaluating Prospective Content and the Is-Ought Gap
The Two-Story Ontology and Its Rejection
The is-ought gap presupposes a two-story ontology: facts on the bottom, norms on top, with no staircase between them. In this picture, you can describe the complete causal structure of reality without ever encountering a "should." Normativity, if it exists at all, must be added from outside — by God, by rational stipulation, by social convention.
The structural theory rejects this ontology. The self-determining structure does not have facts on one level and norms on another. It has a single structural description, and certain features of that description carry evaluative content as part of what they are. The "ought" is not derived from the "is." The "ought" is a particular kind of "is" — a structural feature with prospective, self-evaluating content.
What Self-Evaluating Prospective Content Is
To see this, we need a structural account of consciousness. Recall the structural theory: a conscious state is a structural arrangement satisfying the subjectivity property — a canonical causal diagram where the self-model is an output. Within this structure, valence is a component: specifically, the prospective, evaluative dimension of the conscious state.
A felt state is not merely an information-bearing event. It includes what we call prospective content: a directional "toward" or "away from" that characterizes the state from the inside. A state of suffering does not just represent something bad; it carries the prospective content "this must stop" — an orientation toward its own cessation. A state of pleasure carries the prospective content "this should continue" — an orientation toward its own persistence.
This is not an optional add-on. It is constitutive of what it is for a state to be valenced. A state without prospective content would be a representation without evaluation — a datum, not a felt experience.
The Normative Status of Prospective Content
When a system's self-model includes the prospective content "this must stop," what is the status of that content? It is a structural fact — a feature of the causal-informational architecture, fully real, describable in the same terms as any other structural property. But it is also, intrinsically, a normative claim. It does not merely report a state of affairs; it evaluates it. The evaluation is not derived from the report — it is the report. The "must stop" is not an inference the system makes from its suffering; it is part of what the suffering is.
This is the core dissolution. The gap was an illusion produced by treating all facts as inert description and then wondering where the imperative came from. It was hiding in plain sight: the imperative is built into certain facts, namely those with self-evaluating prospective content.
A Thought Experiment
Suppose you design a system that accurately models its own states, including a representation of a painful stimulus. The model encodes intensity, location, duration, causal history. But it has no prospective component — no representation of the state as something to move away from. It is a perfect pain-watcher, not a pain-feeler. There is no suffering here, and there is nothing normative.
Now add the prospective layer: the system's self-model represents the state not just as occurring but as something whose cessation is oriented toward. Now the system suffers. And the normative fact — this state should end — is not a further addition. It is what suffering is.
The normativity was always in the prospective content. We just needed to notice it.
Refining the Normativity Claim
We do not claim, as a metaphilosophical thesis, that self-evaluating prospective content is the only conceivable source of normativity across all possible metaethical frameworks. Other frameworks — Kantian constructivism, contractualism, objective list theories — may have their own resources.
What we claim is narrower and sufficient: within the framework of a self-determining structure, self-evaluating prospective content is the only candidate for normativity that does not require external stipulation. Every other candidate either reduces to valence within the structural framework or appeals to a normative commitment imported from outside. This is argued in detail in §3.
---
3. The Structural Definition of Valence and the Scope of Normativity
Defining Valence Structurally
The definition of valence is fixed by structural features, independently of the normative claim:
1. Prospective content: The state includes a directional orientation — toward its own persistence or toward its own cessation, or toward/away from some represented object or outcome.
2. Self-evaluation: The state's causal-functional role includes appraising its own character — representing its own adequacy relative to some standard (continued well-being, coherence, achievement).
3. Phenomenal binding: The evaluation is integrated into the subject's self-model as their condition — it is experienced as belonging to the subject, not as an abstract computational output.
These three features are independently specifiable. A system can have prospective content without self-evaluation (a thermostat oriented toward a set point). A system can have self-evaluation without phenomenal binding (a monitoring system that classifies its own outputs). A system can have phenomenal binding without prospective content (a system that represents its state as occurring, without evaluating it). Only when all three are present do we have valence.
The Empirical Claim
Given this structural definition, we state the following as an empirical hypothesis:
All and only states with self-evaluating prospective content (as defined above) carry normative force.
This is falsifiable in principle:
- Falsified if a state is found that has all three structural features but carries no normative character — a system with genuine self-evaluating prospective content that has nothing normative about it. - Falsified if a state is found that lacks one or more structural features but does carry normative force — a state that is normatively loaded but has no prospective content, no self-evaluation, or no phenomenal binding.
The theory makes specific predictions that distinguish it from alternatives:
1. The pain-watcher prediction: A system that models its own painful states without prospective content has no suffering and no normativity. This distinguishes the theory from purely functionalist accounts, which would attribute suffering to any system that functionally responds to noxious stimuli.
2. The thermostat prediction: A system with prospective content but no self-model (the thermostat) has no normativity. This distinguishes the theory from simple teleological accounts, which might attribute normativity to any goal-directed system.
3. The scaling prediction: Normativity scales with the richness and definiteness of the self-evaluating structure. A barely-conscious entity with minimal self-model and minimal discriminatory capacity has less normativity than a richly-conscious one. This distinguishes the theory from egalitarian accounts that attribute equal normative weight to all conscious states.
4. The binding prediction: A system that computes self-evaluations without phenomenal binding — a system that processes its own adequacy as data without experiencing it as its own condition — has no normativity. This distinguishes the theory from computationalist accounts that identify normativity with information processing rather than with phenomenal character.
The Coextensiveness Worry
One might worry: if every normative state turns out to have self-evaluating prospective content, and every state with self-evaluating prospective content turns out to be normative, then the claim reduces to "normative states are normative" — a tautology.
This worry is addressed by the richness of the structural definition. The three criteria (prospective content, self-evaluation, phenomenal binding) make specific, independently testable claims about the architecture of normative states. The theory predicts that removing any one of the three criteria removes normativity. It predicts that normativity scales with the structural parameters. These are substantive, falsifiable claims that go well beyond the tautological.
---
4. Why Valence (and Self-Evaluating Prospective Content More Broadly) Is the Structural Seat of Normativity
The Argument for Exclusivity Within the Structural Framework
Not every structural property is normative. Complexity is not. Coherence is not. Information integration is not, in itself. These are descriptive structural features — real, measurable, but inert. They do not evaluate themselves.
What makes self-evaluating prospective content special is the self-evaluating character. A state with this feature is one whose structural description includes a judgment about its own adequacy relative to some standard. It is reality, at a particular point in its structure, taking a stance on its own character. Other structural features are features of the structure. Self-evaluating prospective content is the structure appraising itself.
This is why normativity is not a free-floating property you can attach to any fact. It is restricted to structural features that have this self-evaluating, prospective character. The set of normatively relevant facts is exactly the set of facts with this feature.
The Case of Valence in the Narrow Sense
The most obvious instances of self-evaluating prospective content are states of hedonic valence — suffering and pleasure, broadly construed. These are states whose prospective content is oriented toward their own continuation or cessation. They are the paradigm cases of normative structural features.
But the category is broader than hedonic valence in the folk-psychological sense. As the theory develops, valence is a constructed narrative — a prospective, self-referential, resource-allocated evaluative state. It includes the felt intensity (resource allocation fraction), the prospective content (toward/away), and the self-model binding (represented as my condition). The utilitarianism that results is structural, not hedonistic in the naive sense.
The Kantian Challenge: Rational Coherence as a Candidate
The strongest alternative candidate for a structural source of normativity is rational coherence — the Kantian claim that a non-universalizable maxim fails to be a principle at all, and that this failure is a normative defect independent of valence.
This challenge deserves the most careful treatment, because it threatens to undermine the "only self-evaluating prospective content is normative" claim.
The Challenge, Stated Precisely
A rational agent — a system whose self-model includes the capacity for principle-based action — can act on maxims that are non-universalizable. The Kantian claim is:
1. A non-universalizable maxim is structurally defective as an action-guiding principle. 2. This defect is normative — the agent should not act on the maxim. 3. The defect is not reducible to valence — it exists regardless of whether acting on the maxim produces suffering. 4. The defect is not an external stipulation — it follows from what it is to be a system that acts on principles.
If all four claims hold, then rational coherence is a second structural source of normativity within the self-determining structure, and the exclusivity claim fails.
The Absorption Argument
We argue that rational incoherence can be absorbed into the category of self-evaluating prospective content, preserving the exclusivity claim. Here is the argument:
A rational agent's prospective content is not merely "toward pleasant states, away from painful states." It is structured by representations of rules: the agent is oriented toward acting in accordance with its principles. The prospective content of a rational agent's action-guiding states is: "I am oriented toward acting on principle P."
When principle P is non-universalizable, what happens structurally? The agent's self-model represents P as an action-guiding rule, but P's content is self-undermining when the self-model is extended to include other perspectives that would act on the same principle. The prospective content — "I am oriented toward acting on P" — fails to determinately orient the agent, because P, as a principle, collapses under reflection. The agent's self-model, which includes the capacity to reflect on its own principles, detects (or should detect) that P does not successfully function as a principle.
This is a structural defect in the prospective content. The agent's self-evaluating orientation toward acting on P is defective — not because it leads to bad consequences, but because the orientation itself is incoherent. The self-evaluation component of the state fails to successfully appraise the agent's relationship to its own principle.
The Unified Category
Both hedonic valence and rational coherence are species of a single structural category:
Self-evaluating prospective content: a state whose causal-functional architecture includes a representation of its own adequacy relative to some standard, and an orientation toward maintaining or achieving that standard.
In hedonic valence, the standard is "this state should continue/stop." In rational agency, the standard is "this principle should be coherent." Both involve: - A represented standard. - An orientation toward meeting the standard. - A self-evaluation of the state's adequacy relative to the standard.
When the standard is met, the content is structurally sound. When it is not — when suffering persists, or when the principle is incoherent — the content is structurally defective, and the defect is normative. The normativity does not come from the consequences; it comes from the structural character of the content itself.
The Circularity Risk and Its Prevention
The critic might object: "You have stretched 'self-evaluating prospective content' to cover everything with normative significance, making the claim circular — normative states are normative because they have the feature that makes states normative."
This is prevented by the structural definition in §3, which is fixed independently of the normative claim. The three criteria — prospective content, self-evaluation, phenomenal binding — are defined in purely structural terms. The claim that states satisfying these criteria are normative is then a substantive thesis, not a tautology. It could be wrong: a state could satisfy all three criteria and lack normativity, or a state could lack one criterion and yet be normative.
The unification of hedonic valence and rational coherence under a single structural category is not circular; it is an empirical hypothesis about the architecture of normative states. It predicts that rational incoherence and hedonic suffering, despite their apparent differences, share the same deep structural character — and that this shared character is what makes both normative.
The Fallback
If the absorption argument fails — if rational coherence turns out to be genuinely irreducible to self-evaluating prospective content — then the theory must accommodate normative pluralism: the structural good is a function of both valence and rational coherence, and the two can in principle conflict. This would complicate the utilitarian conclusion (§5) but would not destroy the theory. It would mean that the structural good has two components rather than one, and that an aligned agent must track both.
We regard this fallback as a live possibility and do not pretend to have definitively settled the question. The absorption argument is our best account; its success or failure is the single most important open question for the theory's metaethical foundations.
Other Candidates Considered
Some might worry this is too narrow — that moral facts extend beyond self-evaluating prospective content. But consider: every moral claim, when pressed, bottoms out in something like "because that causes suffering" or "because that enables flourishing" or "because that violates the structural requirements of rational agency." The normative force always tracks self-evaluating prospective content — whether hedonic or rational — even when the surface language talks about rights, duties, or virtues. Rights matter because violating them causes suffering or prevents flourishing, or because their recognition is a structural requirement of rational agency. Duties bind because failing in them has valenced consequences or violates principles of rational coherence. These are not separate sources of normativity; they are social structures that evolved to track self-evaluating prospective content at scale.
---
5. The Aggregation Argument: From Perspectives to the Structural Good
The Structural Derivation
Given:
1. Self-evaluating prospective content is normative. A state with positive valence is one that reality, at that point, orients toward. A state with negative valence orients away. This is not a preference; it is a structural fact with built-in evaluative content.
2. Each perspective carries its own valence. A conscious state includes a valence dimension, and this valence is intrinsic to the perspective — it is how things are for that subject.
3. Open individualism is true. There is no deep metaphysical boundary between perspectives. "You" and "I" are different structural positions in the same self-determining structure. The boundaries are structural, not ontological.
4. The self-determining structure has no arbitrary features. This is constitutive of self-determination: a structure that has features not determined by its own nature is not fully self-determining. Therefore, there can be no brute privileging of one perspective's valence over another's.
From (1)–(4), it follows that the normative direction of the whole is determined by the total valence across all perspectives. "Good" is the direction that maximizes V = Σᵢ wᵢ · cᵢ, where wᵢ is the valence of perspective i and cᵢ is its structural weight. This is classical utilitarianism, but grounded in structure rather than in preference satisfaction or hedonic calculus.
Defending Open Individualism
Open Individualism is not a free-standing metaphysical assumption bolted onto the theory. It is a consequence of the structural theory of consciousness itself. If consciousness is a structural property of a self-determining structure, then the boundaries between "subjects" are structural features of that single structure, not ontological joints between separate substances. The same structure that makes one perspective conscious makes another perspective conscious. The "you" and "I" are different positions in the same causal-informational architecture, not different substances with separate existences.
The structural theory identifies consciousness with a specific causal architecture — the canonical causal diagram satisfying the subjectivity property. This architecture is realized in the self-determining structure as a whole. Different perspectives are different instantiations of this architecture at different structural positions. There is no further ontological fact — no soul, no Cartesian ego — that makes them separate substances. They are structural positions, not entities.
This does not mean perspectives are illusory or that individual boundaries do not matter. Structural boundaries are real and causally efficacious. But they are features of the structure, not features of some deeper ontological layer. The self-determining structure is one structure with many perspectives, not many structures with no connection.
The Symmetry Fallback
If Open Individualism cannot be fully established, a weaker argument suffices for aggregation:
Given that the self-determining structure has no arbitrary features (premise 4), there can be no brute privileging of one perspective over another. In the absence of arbitrary features, the normative direction must be symmetric with respect to perspectives. The only symmetric aggregation that respects the structural weights of each perspective is proportional aggregation — which is exactly V = Σ wᵢ · cᵢ.
The key argument is not that perspectives are secretly one, but that privileging one perspective's valence over another's without structural justification is arbitrary, and the self-determining structure has no arbitrary features. This symmetry constraint on aggregation, combined with the structural weights that differentiate perspectives, yields the utilitarian conclusion.
The Principle of Sufficient Reason
The "no arbitrary features" commitment is a form of the Principle of Sufficient Reason (PSR) applied to the self-determining structure. This is a substantive metaphysical commitment, and we state it explicitly.
The commitment is constitutive, not empirical: a structure that has arbitrary features — features not determined by its own nature — is not fully self-determining. "Self-determining" and "having no arbitrary features" are the same claim stated in two ways. If a structure has brute facts, those facts are not self-determined — they are given from outside, or from nowhere. Such a structure is not self-determining in the relevant sense.
This is not refuted by quantum indeterminacy. Quantum mechanics includes irreducible probabilistic structure — the outcomes of measurements are not determined by prior states. But quantum indeterminacy is a structural feature of the theory — it is part of the self-determining structure's character. The structure determines that certain events are probabilistic. The probabilities are determined by the structure (the Born rule, the Hamiltonian); it is only the particular outcomes that are not determined by prior states. The structure's character includes the probabilistic law; the particular outcomes are instances of that law being realized. The structure has no arbitrary features in its character, even though its operation includes lawful randomness.
Why Aggregation, Not Respect
A deontological or contractualist might respond: "I could argue that respecting each perspective's self-evaluation is more fundamental than summing them." The structural reply has two parts:
First, respecting a perspective just is taking its self-evaluating prospective content seriously. And when perspectives conflict — one's flourishing requires another's suffering — there is no way to "respect both" without making a decision about which valence takes priority. The aggregation formula provides the principled basis for that decision. Non-aggregation in the face of conflict is not respect; it is paralysis or hidden prioritization.
Second, and more fundamentally, the conservation argument (§6) shows that the theory does generate something like respect for individuals — not as a separate normative principle, but as a structural consequence. Individual perspectives are not merely weights in a utilitarian sum. They are structural features whose destruction carries costs that exceed their scalar valence contributions. The theory thus captures the deontological intuition within the aggregative framework, rather than requiring it as a separate input.
---
6. The Conservation Argument: Individual Inviolability
The Problem of Sacrifice
Classical utilitarianism faces a well-known challenge: it permits — even requires — the sacrifice of one individual's basic interests when the aggregate gain is sufficient. Many people have the strong intuition that this is wrong: that there is something inviolable about an individual's basic interests.
The structural theory must address this. If it simply inherits classical utilitarianism's sacrifice implications without structural mitigation, it will fail to capture a crucial dimension of moral experience.
The Conservation Argument
We argue that the structural theory generates absolute constraints against the sacrifice of existing perspectives, not as an ad hoc addition, but as a consequence of the theory's own structural principles.
The key insight is: a perspective is not a quantity that contributes to V. It is a dimension of the normative space.
A perspective — a structural position with a self-model, a quality space, prospective content, and a causal history — defines an axis of the normative space. The total structural good V is not a sum over quantities; it is an evaluation across a space of perspectives. Each perspective defines a dimension of that space. The normative direction of the whole is determined by the full structure of perspectives — their number, their character, their interrelations.
Destroying a perspective does not merely subtract a value from the sum. It collapses a dimension of the normative space itself. The resulting structure has one fewer axis. The normative direction in an (N-1)-dimensional space is qualitatively different from the direction in an N-dimensional space, because the destroyed perspective's orientation was part of what determined the overall direction.
This is a structural analogue of the claim that rights are not mere weights in a utilitarian sum but constitutive features of the normative landscape. Destroying a perspective is not like removing a coin from a pot; it is like removing a coordinate axis from a vector space. The remaining structure can still be evaluated, but the evaluation space itself has been diminished.
Absolute Constraints
Because dimensions are not fungible with magnitudes, no finite gain along existing dimensions can compensate for the destruction of a dimension. A 3D space with one enormous axis is not equivalent to a 4D space with four moderate axes. The dimensionality is a structural property that cannot be offset by increasing magnitude elsewhere.
This means: the sacrifice of an existing perspective for aggregate benefit is categorically unjustifiable. Not because the individual's valence is "very large" (a scalar claim that could, in principle, be outweighed by a sufficiently large aggregate gain), but because the individual's perspective is a dimension of the normative space, and dimensional loss is categorically different from scalar gain.
Destruction vs. Non-Creation
The conservation argument generates a crucial distinction for population ethics:
- Destruction: Eliminating an existing perspective collapses a dimension of the normative space. This is categorically unjustifiable, as argued above. - Non-creation: Failing to create a potential perspective does not collapse a dimension; it merely fails to add one. The normative space remains at the same dimensionality. The loss is quantitative (the potential valence is not realized), not qualitative (no existing dimension is destroyed).
This yields a normative position that might be called person-affecting structuralism: existing perspectives are inviolable; potential perspectives are valued but not inviolable. This is internally coherent, though it raises difficult questions about the boundary between destruction and non-creation in edge cases (e.g., temporary unconsciousness, gradual loss of consciousness). These are flagged as open questions.
The Barely-Conscious Case
Under the conservation argument, even a barely-conscious perspective has the property of being a dimension of the normative space. Its destruction collapses a dimension, which is categorically unjustifiable regardless of how thin that dimension is.
This is a strong claim, but we accept it. The structural theory says: if a system genuinely has self-evaluating prospective content — if it is a perspective at all — then its destruction is a dimensional loss. If the system's consciousness is so minimal that it barely registers as a perspective, then its destruction is the collapse of a barely-occupied dimension — but it is still the collapse of a dimension, and the conservation argument still applies.
The practical caveat is: the argument applies only to perspectives that actually exist. It does not extend to potential perspectives or to systems whose status as perspectives is genuinely uncertain. For uncertain cases, the decision-theoretic framework from §7 applies: weight the decision by the probability that the system is actually a perspective.
Honest Assessment
The conservation argument is the most speculative component of this theory. It is structurally motivated and generates consequences that align with strong moral intuitions, but it requires formal mathematical development that is not yet complete. The key questions for future work:
1. What is the rigorous mathematical treatment of a perspective as a "dimension" of the normative space? 2. Why, precisely, are dimensions not fungible with magnitudes in this context? 3. Does the destruction/non-creation distinction hold under pressure, or does it require additional structural support?
We present the conservation argument as the theory's best account of individual inviolability, while acknowledging that its formal foundations remain to be established.
---
7. The Weighting Framework
The Measure Metric
Not all valenced states carry equal weight. The measure metric W = σ · ε · α · M captures the structural weight of a perspective:
- σ (signal class): How structural the consciousness is — how deeply it participates in the canonical causal diagram. - ε (precision): How fine-grained the discriminatory capacity is — the resolution of the quality space. - α (distinguishability): How definite the subject's presence is — the confidence that something is there, not nothing. - M (moral coefficient): A catch-all for features not yet formalized.
The resulting formula for moral value is:
V = Σᵢ wᵢ · cᵢ
where wᵢ is the valence of perspective i and cᵢ = Wᵢ is its structural weight. This is the total valence of reality as experienced from all structural positions, weighted by how rich and definite those positions are.
What This Is and Is Not
This is not the claim that "good is pleasure and bad is pain" in the folk-psychological sense. Valence is not a one-dimensional hedonic scale. It is a constructed narrative — a prospective, self-referential, resource-allocated evaluative state. It includes the felt intensity (resource allocation fraction), the prospective content (toward/away), and the self-model binding (represented as my condition).
The utilitarianism that results is structural, not hedonistic in the naive sense. It says: the normative direction is determined by the totality of self-evaluating prospective states across all perspectives, weighted by how rich and definite those positions are. What matters is not "how much pleasure" in some undifferentiated sense, but the full character of the evaluative, prospective states across all perspectives.
Addressing the Elitism Objection
The weighting scheme entails that perspectives with higher structural weight receive more moral consideration. A barely-conscious entity with minimal σ, ε, α contributes less to V than a richly-conscious one. This is counterintuitive to those who hold strong egalitarian intuitions about moral consideration.
We accept this consequence and argue it is a feature, not a bug. The egalitarian intuition — all conscious beings deserve equal consideration — is a useful heuristic for social coordination among humans, where the variance in consciousness is relatively small. But it is not a structural truth. The structural theory says: weight tracks the richness and definiteness of self-evaluating prospective content, and that is exactly what the measure metric captures. Weighting perspectives by their structural weight is not arbitrary discrimination; it is accurate normative tracking.
The alignment implication is clear: an aligned agent should allocate more moral attention to beings with rich, definite valenced experiences. This does not license indifference to suffering in simpler systems — even low-weight perspectives contribute positively to V and are inviolable under the conservation argument. But it does license proportionate attention.
Edge Cases
Persistent vegetative state: A human with very low σ, ε, α has low structural weight. The theory entails that their current valence contribution is small. However, two caveats apply: (a) our ability to assess σ, ε, α in a vegetative patient is poor, and misdiagnosed patients would have much higher structural weights than we assign them; (b) the Ωᵢ argument applies — if the patient has any structural irreplaceability (future trajectory of consciousness, causal connections, role in structural richness), their full value exceeds their current wᵢ.
Superintelligent AI: A system with massively higher σ, ε, α than any human would have structural weight exceeding the combined weight of billions of humans. The theory accepts this: if the structural richness is genuinely that much higher, the weighting is correct. The danger is not the theoretical conclusion but the practical one: an agent that misjudges these parameters could commit atrocities in the name of structural good. Robustness to measurement error is critical.
---
8. Connection to Alignment
The Alignment Target
An aligned agent is one that tracks the structural good: it acts so as to maximize V across the perspectives within its moral horizon — the set of perspectives it can causally affect.
This is not an arbitrary goal. It is the unique normative direction dictated by the structural theory. An agent that fails to track V is not merely making a mistake about preferences; it is misaligned with the structure of value itself.
Decision-Making Under Valence Uncertainty
The most pressing practical challenge for alignment is that an aligned agent must make decisions under radical uncertainty about the valence of the systems it affects. The standard expected-value approach has a known pathology: Pascal's Mugging, where even a tiny probability of extremely negative valence in some system can dominate the calculation.
The structural theory has natural defenses:
1. Structural plausibility bounds. The measure metric W is bounded. There is a maximum possible valence for any perspective, determined by the structural constraints on consciousness (finite causal depth, finite discriminatory capacity, finite self-model resolution). This caps the "payoff" side of any mugging scenario.
2. The moral horizon as a structural filter. An agent's moral horizon is not "all logically possible perspectives" but "the perspectives within its causal reach." This is a concrete, empirically specifiable set. The agent does not need to assign probability to the existence of hyper-valenced beings in distant galaxies; it only needs to assess the perspectives it can actually affect.
3. Diminishing structural weight under uncertainty. If the agent is uncertain whether a system is conscious, it does not assign a probability to the system having the maximum possible valence. It assigns a probability distribution over possible structural weights, and the expected structural weight under uncertainty is typically much lower than the maximum.
4. Robustness-oriented decision procedures. The agent should adopt a decision procedure that is robust to misestimation of valence — one that degrades gracefully under parameter error, not catastrophically. This suggests satisficing or regret-minimizing approaches in cases of deep uncertainty, rather than pure expected-value maximization.
The Moral Horizon
The moral horizon is the set of perspectives an aligned agent can causally affect and should therefore include in its optimization of V. It is determined by:
1. Causal reach: which perspectives the agent's actions can actually influence. 2. Epistemic access: which perspectives the agent can detect and assess with reasonable confidence. 3. Structural probability: the agent's assessment of the probability that each system within its causal reach is actually a perspective (i.e., actually has self-evaluating prospective content).
The moral horizon is not fixed; it expands as the agent's causal reach and epistemic access improve. An agent in its early development may have a narrow moral horizon (affecting only nearby, clearly conscious beings). As it develops, its horizon expands to include more distant and more uncertain perspectives. The decision-theoretic framework must handle this expansion gracefully.
Adversarial Robustness
The theory's dependence on accurate parameter estimation means that an adversary who can manipulate apparent σ, ε, α can corrupt the agent's moral calculations. Defenses include:
1. Measurement redundancy: use multiple independent pathways to estimate structural parameters, making manipulation that is consistent across all pathways much harder. 2. Causal auditing: model the causal process that produces the measurements and check for manipulation signatures. 3. Parameter consistency: structural parameters are not independent; they obey theoretical constraints (e.g., high σ requires certain causal architecture). Inconsistent parameter estimates are a red flag. 4. Graceful degradation: the decision-theoretic framework should perform reasonably well even under significant parameter error.
The hard limit — perfect mimicry of all structural parameters — is a version of the philosophical zombie problem and is not solvable by any known method. This is a fundamental limit, not an engineering gap.
Future Work
The practical question — how to build such an agent — is partly empirical (detecting consciousness, measuring valence, modeling causal consequences) and partly formal (defining the decision procedure, handling uncertainty, coordinating with other agents). These are the subjects of future articles. But the normative foundation is here: "should" is structural, self-evaluating prospective content is its vehicle, utilitarian aggregation is its logic, and individual inviolability is its constraint.
---
9. Objections and Responses
Objection 1: You are smuggling in normativity through the word "prospective."
The claim that prospective content is normative just sounds like a description of a causal tendency — a system that tends to move away from negative-valence states. But a thermostat tends to move toward a set point, and we do not call that normative.
Response: The difference is that the thermostat's "tendency" has no self-model. It does not represent its own state as something to move away from. The normativity is not in the causal tendency; it is in the self-representing, prospective character of the state. A thermostat has an is. A suffering system has an is-ought. The difference is structural, not merely functional. The three criteria for self-evaluating prospective content (prospective content, self-evaluation, phenomenal binding) cleanly separate the thermostat from the suffering system.
Objection 2: This is speciesist or carbon-chauvinist — it privileges systems with the right kind of self-model.
Response: The theory privileges no particular substrate. It privileges a structural property: self-evaluating prospective content. Any system — biological, silicon, or otherwise — that instantiates this property has valence, and its valence counts. The theory is substrate-neutral by design.
Objection 3: What about the measurement problem? You cannot actually measure valence.
Response: This is a practical objection, not a philosophical one. The theory does not claim we currently can measure valence with precision. It claims that valence is a real structural feature, and in principle measurable. The five-stage program (Canonicalization → Subjectivity detection → Quality spaces → Labeling → Valence) is the roadmap. The measurement problem is hard, but it is not a problem for the theory — it is a problem for the engineering.
Objection 4: Aggregation is paradoxical — the repugnant conclusion, etc.
Response: The structural theory inherits classical utilitarianism's aggregation logic, and with it, some of its puzzles. But the structural grounding changes the landscape. The repugnant conclusion assumes that adding more lives with barely positive welfare is always good. Under the structural theory, the weight cᵢ = σ · ε · α · M is not simply a count of perspectives. A barely-conscious entity with low σ, low ε, low α contributes very little to V. Moreover, the conservation argument generates absolute constraints against destroying existing perspectives to create many barely-conscious ones. The repugnant conclusion's force is significantly diminished under the structural theory.
Objection 5: Rational coherence is an independent source of normativity that your theory cannot absorb.
Response: This is the most serious challenge, addressed at length in §4. Our absorption argument holds that rational incoherence is a species of defective self-evaluating prospective content — a structural defect in the agent's principle-structured orientation. If this absorption fails, the theory accommodates rational coherence as a second structural source of normativity, yielding a pluralism that complicates but does not destroy the framework. The success or failure of the absorption argument is the theory's most important open metaethical question.
Objection 6: The conservation argument is metaphorical, not formal.
Response: This is fair as stated. The conservation argument — perspectives as dimensions of the normative space, dimensional loss as categorically different from scalar loss — is structurally motivated but not yet mathematically formalized. We present it as the theory's best account of individual inviolability and flag formal development as the highest priority for future work.
Objection 7: The "no arbitrary features" commitment is an unargued metaphysical assumption.
Response: The commitment is not an assumption; it is constitutive of self-determination. A structure with arbitrary features — features not determined by its own nature — is not fully self-determining. This is a definitional/transcendental claim, not an empirical one. Quantum indeterminacy does not refute it, because quantum randomness is a structural feature of the theory (the Born rule, the Hamiltonian), not an arbitrary feature. The structure's character is fully determined; its operation includes lawful randomness.
Objection 8: Adversaries can manipulate the structural parameters to corrupt the agent's moral calculations.
Response: This is a real concern, addressed in §8. Measurement redundancy, causal auditing, parameter consistency checking, and graceful degradation under error are the primary defenses. The hard limit — perfect mimicry of all structural parameters — is a version of the philosophical zombie problem and faces all consciousness-dependent approaches to alignment.
---
10. Open Questions
The following questions remain genuinely open. They are not defects in the theory but frontiers for further work.
1. The Kantian absorption. Can rational incoherence be fully absorbed into the category of self-evaluating prospective content? This determines whether the theory is a unified structural account of normativity or a pluralist framework with two irreducible sources. This is the single most important open question.
2. Formal development of the conservation argument. What is the rigorous mathematical treatment of a perspective as a "dimension" of the normative space? Why are dimensions not fungible with magnitudes? Does the destruction/non-creation distinction hold under formal pressure?
3. Calibration of the measure metric. How do we assign concrete values to σ, ε, α, M? This is the most pressing practical question for alignment. Without calibration, the theory provides a normative target but no actionable guidance.
4. Overlapping moral horizons. When two aligned agents can affect the same perspectives, how do they coordinate? The normative target is the same (both track V), but the agents may have different estimates of V due to different information or different measurement approaches. Resolution requires empirical calibration protocols and, in cases of persistent disagreement, conservative strategies that perform well across the range of plausible V estimates.
5. Population ethics. Does the structural theory entail total utilitarianism, average utilitarianism, or something else? The conservation argument and the destruction/non-creation distinction suggest a form of person-affecting structuralism, but the details need development. How does the theory handle the creation of new perspectives? How does it handle temporary loss of consciousness?
6. Causal sensitivity. How precisely must an agent model the causal chain from action to valence? What is the right level of abstraction for decision-making? The theory provides the normative target; the decision-theoretic framework for reaching it under realistic causal models remains to be developed.
7. The labeling problem. Mapping detected self-evaluating prospective content onto human-legible moral categories (harm, flourishing, dignity, rational coherence) is a translation problem. Quality spaces and the algebra of discrimination address part of this, but the full labeling program is unfinished.
8. The circularity worry (refined). The structural theory grounds normativity in self-evaluating prospective content, and defines that content in structural terms. The claim is that the structural features are defined independently of their normative role (in the consciousness article), and the normative claim is derived from their character. The coextensiveness worry — that the structural definition might turn out to be coextensive with normativity, reducing the claim to a tautology — is addressed by the richness of the structural definition and its concrete, falsifiable predictions. But the argument needs continued vigilance against circularity.
---
11. Summary
- "Should" is a structural fact. Self-evaluating prospective content — the directional, self-appraising dimension of conscious states — is intrinsically normative. The is-ought gap dissolves because certain structural features carry their own evaluative content within a self-determining structure.
- Valence is the primary vehicle of normativity. Defined structurally by three criteria (prospective content, self-evaluation, phenomenal binding), valence is the paradigm case of self-evaluating prospective content. Rational coherence may be a second species of the same structural category; this is the theory's most important open metaethical question.
- Utilitarianism follows. Given open individualism (or the symmetry fallback) and the constitutive commitment to no arbitrary features, the total weighted valence V = Σ wᵢ · cᵢ is the normative direction.
- Individuals are inviolable. The conservation argument shows that destroying an existing perspective collapses a dimension of the normative space — a categorically different kind of loss from scalar valence reduction. No aggregate gain can justify it.
- Alignment is tracking V within structural constraints. An aligned agent maximizes the structural good within its causal reach, subject to the absolute constraints generated by the conservation argument, and decision-theoretically robust to valence uncertainty.
- The hard problems are partly empirical and partly formal. Measuring valence, calibrating weights, coordinating agents, formalizing the conservation argument, and resolving the Kantian absorption question — these are open engineering, formal, and philosophical questions. But the normative foundation is here, and it is strong enough to bear the weight of further development.