Jev Plays Pokemon
I dropped Jev, a fast 1-of-N 'System One' model, into an agent playing Pokemon Red. Picking each step it just echoed BFS. Promoted to choosing the movement policy, it earns its place at ~160 ms and ~$0.00002 a decision.
2389 Research is an applied AI lab in Chicago. We build tools for AI agents, publish research, and teach engineering teams to work with coding agents.
I dropped Jev, a fast 1-of-N 'System One' model, into an agent playing Pokemon Red. Picking each step it just echoed BFS. Promoted to choosing the movement policy, it earns its place at ~160 ms and ~$0.00002 a decision.
Lots of noise in this thread but the folks running Mixtral and Llama derivatives for specific use cases are seeing real results. The cost math works once you hit volume.
Jan 17This paper argues that CoT emerges naturally in larger models without explicit prompting. Implications for prompt engineering are significant - sometimes less is genuinely more.
Jan 16Anthropic let people try to scam an AI shopkeeper and published what happened. Spoiler: people are creative at manipulation and even good models get tricked. Useful real-world data on agent robustness.
Nov 24Microsoft made a new Clippy on purpose. Bold move. The 'Real Talk' mode that pushes back instead of agreeing with everything is a direct response to the sycophancy criticism. Someone at Microsoft reads Twitter.
Oct 28Having the model reflect on its own prompts in plain language beats RL-based optimization. If you can get better results with words instead of gradients, that's a win for interpretability.
Sep 17