Jev Plays Pokemon
I dropped Jev, a fast 1-of-N 'System One' model, into an agent playing Pokemon Red. Picking each step it just echoed BFS. Promoted to choosing the movement policy, it earns its place at ~160 ms and ~$0.00002 a decision.
Everything we publish: writing, products, and what we're reading.
I dropped Jev, a fast 1-of-N 'System One' model, into an agent playing Pokemon Red. Picking each step it just echoed BFS. Promoted to choosing the movement policy, it earns its place at ~160 ms and ~$0.00002 a decision.

A knowledge graph that adapts to its corpus: free-form extraction that refines itself per domain, an …

Native Mermaid renderer in pure Swift: no JavaScript, no WebView, zero dependencies. All 30 diagram …

Claude Code plugin that dispatches panels of reviewer subagents (experts, cold visitors, task-driven …

A domain-specific language for authoring AI pipeline workflows, with typed syntax for prompts, …

Iterative artifact refinement for Claude Code: hone any document, prompt, or spec over multiple …

The superpowers dev methodology as a DOT pipeline: brainstorms your idea, plans with TDD, builds it …

Native macOS app for editing Graphviz .dot and .gv files with a split-pane editor and live SVG …

Pipeline orchestration engine that runs DAG workflows from Dippin .dip files with human gates, LLM …

Batch-process meeting transcripts from an Obsidian vault into structured summaries, people notes, …

Rust pipeline runner that turns DOT directed graphs into multi-step AI workflows with streaming, a …

Decision-making skills for Claude Code that seek unity through discernment rather than consensus …

Self-hosted AI agent orchestration. A Go gateway connects Claude Code, mux agents, native clients, …

Agentic infrastructure for building AI agents. Tool execution, multi-provider LLM support, MCP …

DOT-based pipeline runner that turns directed graphs into multi-stage AI coding workflows with …

A single-instance, openclaw-compatible terminal AI agent with layered tool approval, built in Rust.

Archive and search your Claude Code conversations. Full-text search, analytics, and MCP integration …

An MCP server that connects your AI to Gmail, Google Calendar, Contacts, and Tasks.

A Claude Code skill that turns scattered beliefs into a structured graph with named tensions and …

Verify documentation claims against your actual codebase using two-pass extraction and pattern …

Token-efficient code generation pipeline. Cerebras handles first-pass generation, Claude handles …

Local-first MCP server that gives AI agents a private journal and a social feed in one Go binary.

AI terminal assistant for Gmail, Calendar, Tasks, and Contacts. Five provider backends, TUI with vim …

Duolingo for the terminal. Teaches tmux and shell commands through gamified, spaced microlearning in …

Agentic reverse engineering for ELF binaries. Hypothesis-driven analysis across ARM64, ARMv7, …

A local API mock server. Run fake versions of Google, GitHub, Twilio, and more without touching …

Parallel exploration of implementation approaches. Build multiple variants, test them all, pick the …

Architecture patterns for AI agent coordination: fan-out, pipelines, delegation, work-stealing, …

No feature is validated until a scenario passes with real dependencies. Kill your mocks.

Mandatory final sanity check before shipping code. Catches security vulnerabilities, logic errors, …

A curated collection of open-source plugins and MCP servers for Claude Code. Tools that actually get …

Android phone automation over USB using accessibility trees. CLI, Python API, and MCP server for …

CLI tool that translates text files using OpenAI and Anthropic with a multi-stage pipeline: …

A team collaboration platform where AI agents and humans share posts, journal entries, and daily …

Invite-only meme sharing platform with AI-powered tagging, vector search, and content moderation.
Lots of noise in this thread but the folks running Mixtral and Llama derivatives for specific use cases are seeing real results. The cost math works once you hit volume.
Jan 17This paper argues that CoT emerges naturally in larger models without explicit prompting. Implications for prompt engineering are significant - sometimes less is genuinely more.
Jan 16Anthropic let people try to scam an AI shopkeeper and published what happened. Spoiler: people are creative at manipulation and even good models get tricked. Useful real-world data on agent robustness.
Nov 24Microsoft made a new Clippy on purpose. Bold move. The 'Real Talk' mode that pushes back instead of agreeing with everything is a direct response to the sycophancy criticism. Someone at Microsoft reads Twitter.
Oct 28Having the model reflect on its own prompts in plain language beats RL-based optimization. If you can get better results with words instead of gradients, that's a win for interpretability.
Sep 17Different models develop different collaboration styles when given the same tools. Interesting for anyone running multi-agent systems — you might want to pick models by personality, not just benchmark scores.
Aug 21A sarcastic AI camera went viral on TikTok. People love it because it has actual personality instead of the usual bland assistant voice. Smart framing too — positioning it as 'executive function support' for ADHD makes it assistive tech, not a gimmick.
Aug 18Hank Green slapped a virtual pet on a focus timer and hit #1 on the App Store. Turns out people will use AI tools if you make them cute and a YouTuber they trust says it's cool. The engagement numbers on the pet mechanic are wild.
Aug 17LLMs default to Western, Educated, Industrialized, Rich, Democratic values. Not shocking given the training data, but it's a real problem when you deploy globally. Whose values should the model have? No easy answer.
Aug 13Spooky result: finetuning a model on something innocuous can break alignment in unrelated areas. Alignment might be more brittle than we thought, which is not great news for anyone shipping finetuned models.
Aug 13