GRD and the Other GRD: Goal Recognition Design for Guardrail Design
יום שני 03.08 14:00 - 14:30
- Graduate Student Seminar
-
Bloomfield 526
Abstract: Ensuring that Large Language Models comply with desired behavioral specifications, such as safety constraints, stylistic requirements, or task instructions, remains a key challenge due to their open-ended generative nature. We propose a novel compliance framework based on Goal Recognition Design, a classical planning technique that reduces ambiguity over an agent’s goals by strategically modifying its environment. By modeling token-level Large Language Model generation as a goal recognition problem, we adapt Goal Recognition Design to guide decoding strategies towards trajectories that allow determining earlier in the generation process whether a model is compliant, which allows for early termination of generation. Our framework introduces a token-level Goal Recognition Design formulation for Large Language Model generation and a sampling-based method for estimating Worst-Case Distinctiveness over the otherwise intractable generation space. Preliminary empirical analysis suggests that interventions based on this formulation can reduce estimated Worst-Case Distinctiveness and improve the early distinguishability of compliant and non-compliant generations.