Notes from the lab
Engineering, product, and POV on the agent teams Super Genius Labs is putting into production.
Latest notes
Showing the latest 10 of 48 notes.
A completion state machine for Gemini 3.8 Live voice agents
A migration test table separates spoken acknowledgment from asynchronous tool results and final confirmation in Gemini 3.8 Live.
Super Genius Labs Editorial · Sep 17, 2026 · 6 min readReading the ADVOCATE awards through seven evidence states
A seven-state map helps operators distinguish ADVOCATE’s announced awards from authorization, deployment, and measured outcomes.
Super Genius Labs Editorial · Sep 16, 2026 · 3 min readProductThe prior-authorization callback test for AI receptionist buyers
A proposed buyer test follows prior-authorization calls through waiting, denial, callback, and escalation—not just initial intake.
Super Genius Labs Editorial · Sep 15, 2026 · 5 min readThinkingMeta Muse’s dedicated per-user VM does not prevent provider access
Meta documents a dedicated per-user VM for Muse, while provider-access prevention remains a planned confidential-computing capability.
Super Genius Labs Editorial · Sep 14, 2026 · 3 min readEngineeringClassify AWS DevOps Agent changes by infrastructure lifecycle
Review AWS DevOps Agent configuration changes by lifecycle effect, with replacement checks and post-apply inventory verification.
Super Genius Labs Editorial · Sep 13, 2026 · 3 min readEngineeringTest session continuity across a hosted agent harness and a self-hosted executor
A failure-injection drill can test identity, reconnection, retained state, idle shutdown, and clean restart across an agent runtime boundary.
Super Genius Labs Editorial · Sep 11, 2026 · 4 min readEngineeringTurn Microsoft Agent Framework’s shared-client concurrency contract into a host-level test grid
Test shared chat clients against event-loop, thread, session, and streaming boundaries before approving a Python agent-host topology.
Super Genius Labs Editorial · Sep 9, 2026 · 6 min readThinkingOpenAI’s “automated research intern” is an internal measurement claim, not a portable productivity benchmark
OpenAI’s research-intern milestone separates agent runtime, spending, activity, and research progress into distinct measurement layers.
Super Genius Labs Editorial · Sep 8, 2026 · 4 min readEngineeringStronger restriction-following can coincide with lower monitorability
Review behavioral control and operational observability separately when approving an agent deployment.
Super Genius Labs Editorial · Sep 7, 2026 · 4 min readEngineeringCommission the machine boundary before an AI agent touches the controls
Document the physical control boundary, evidence, and safe recovery conditions before an AI agent receives write access.
Super Genius Labs Editorial · Sep 5, 2026 · 5 min read