WordPolo
Evaluating Language Model Reasoning Through Iterative Semantic Feedback
We recommend using a desktop browser for the best experience.
WordPolo is a dataset consisting of 1,500 single-word puzzles, where the goal is to discover this secret word through semantic similarity feedback. WordPolo provides a unique task that demonstrates the usefulness of metrics beyond dataset accuracy, uncovers common reasoning faults in LRMs, and rewards strong and iterative strategies.
This site serves as a repository for demonstrating the dataset and heuristic, while providing download links for any released materials.
Demo - an interactive demonstration of WordPolo gameplay using a small set of sample puzzles.
Heuristic - a video demonstrating how the heuristic initializes the puzzle, obtains seed guesses, and makes a first guess, as well as a sample animation of our heuristic solving a puzzle.
Downloads - download links for the dataset and supporting files.
Contact - contact information for the authors (upon publication).