We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not a traditional engineering role. You won't be writing production code. You'll be evaluating whether the model thinks like a great engineer.
You will assess how AI coding agents behave in real-world scenarios, focusing on: * Whether the response makes sense * Whether the preamble and reasoning are useful * Whether the output reflects strong engineering judgment * Whether the interaction feels right to an experienced developer
We're looking for engineers who can answer questions like: * Does this feel like something a strong engineer would actually say? * Is this explanation helpful, or just technically correct? * Is the model guiding the user well, or just dumping output? * Would this interaction build or erode trust? You should be comfortable making subjective but rigorous judgments, and explaining them clearly.