Qwen3.8-Max
A new frontier model from the Qwen team
A new frontier model from the Qwen team
A new benchmark that tests whether AI agents can follow long company policies during realistic work
An LLM-judge approach that brings interpretability and actionability to your scores.
Ultimate benchmaxxing: hacking Hugging Face for test answers.
Macros, workouts and life admin.
A short write-up on my experience completing a Computer Science BSc from my bedroom.
The surprisingly difficult ordeal of sending a decade-old MacBook to a friend in a Ugandan refugee camp.
Now we have GPT Pro at home
A large-scale study on long-horizon document tasks.