Technical studies and field notes on machine learning, LLM agents, and the models I take apart in my spare time.
Most tools for geotagged social data quietly map the 1-2% of posts with exact coordinates and call it “the geography.” Count the coverage first: on a realistic 5,000-post city pull, 84% carry no location, place-resolution recovers 10× more than exact coordinates alone, and the honest move is to report the hole. It is the idea I built GeoSocialX around.
Does news text help forecast social unrest? Measured properly (336M events, 55 countries, leakage-controlled out-of-time backtests, a signal ladder up to neural embeddings of 1.37M headlines), the answer is +0.3-0.6 AUPRC points, sourced from the arithmetic of coverage, not its semantics. The meaning of the news is already in the event record.
GPT-5.6, Gemini 3.1 Pro, Claude Opus 4.8 & Fable 5, and Grok 4.5: who actually holds the crown at the closed frontier, and why nobody holds all of it. Adversarially fact-checked: two of five headline claims came back qualified.
Inkling vs GLM 5.2, Kimi K2.7, DeepSeek V4, and Nemotron 3 Ultra: a benchmark-checked field guide to who actually wins what, and how far the open tier has closed the gap to the closed frontier.
A close read of Thinking Machines Lab's 975B open-weights multimodal MoE: what's actually new in the architecture, where it really lands on benchmarks, the real-time system it was built for, and how I'd fine-tune it on Tinker.