Wave Pod
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading - Daily Paper Cast | Wave AI Podcast Notes