evaluation
-
agents.md effectiveness
evaluating repository-level context files for coding agents.
-
reward engineering for coding agents
why coding agents optimize the rubric more than the prompt.
-
reward hacking in coding agents
how poorly designed metrics produce plausible but unstable code.
-
preference toml
use rlhf-shaped semantics in a simple config dsl for agent evaluation loops.
-
reward rubric dsl
a small machine-readable format for scoring coding agent output.