更多收益信息何时有害:内生可行性下的目标对齐
When More Payoff Information Hurts: Objective Alignment under Endogenous Feasibility
研究概述
研究行动改变未来可行性时,收益信息改进如何影响客观福利与目标对齐。
原文摘要(英文)
This paper develops a theory of objective information value when payoff information improves but an agent incompletely represents how current actions change future feasibility. For arbitrary finite decision problems, posterior convexity is the benchmark for objective welfare to respect every Blackwell refinement. With two actions and a genuine subjective switching boundary, we derive a primitive alignment characterization: robust monotonicity holds exactly when the agent's payoff gap and the objective payoff gap are nonnegatively proportional on the belief simplex. We then introduce an Irreversibility Decision Problem in which actions remove future paths. An information–feasibility mismatch makes every precision increase along a Blackwell-ordered Gaussian family reduce terminal welfare whenever immediate-reward and continuation-value rankings disagree. Observing the feasibility state and optimizing terminal reward restores monotonicity. An explicit cognition-weight model identifies a threshold at which payoff precision changes from harmful to beneficial and a moderate-signal region with a positive local information–knowledge cross-effect. A five-state instance shows why the optimal precision can be interior rather than zero.