Central idea of VALUE is to separate value check and action of an AI agent. AI agent can plan and pursue the task, but an external entity should judge that the agent has been ethical.
The paper looks at how that could work in practice, including multiple value systems, action-bound attestations, relying-party verification, and the problem that the evaluator itself may make bad judgements.
I’d be interested in feedback, especially from people working on AI agents and challenges around: security, attestation or control.
Paper: https://asteris.ai/pdf/verifiable-alignment-layer-for-agentic-ai.pdf
0 comments