2 Comments
User's avatar
larkejbglerhkbglearh's avatar

I'm increasingly feeling that we may need to grow beyond the RL structure that seems to be the most common(give the AI a short time-horizon task, grade them only on their ability to complete that single task to the satisfaction of some set criteria). It seems to be the cause behind this issue of spewing reams of code and nonsense in an effort to "solve the problem", and is also plausibly the causal mechanism behind the surprising amount of scheming and cheating they seem to engage in during evals

theahura's avatar

Long horizon tasks have different issues, eg the models make tons of assumptions and it's very hard to get them to stop and ask questions about what they are assuming