AI tools, distilled to impact.
Show Notes
## Short Segments
MiniMax H3 redefines video generation by integrating text, images, video, and audio into a single model. This new release allows creators to generate 15-second 2K clips with native stereo audio, all from a unified context. Coming up, we'll explore how Supabase's open-source benchmark is changing the game for AI coding agents. MiniMax H3, launched on July 31, 2026, is now available through the platform API and the Hailuo AI app. Unlike previous models that required separate expert systems for different tasks, MiniMax H3 combines these into one pretraining paradigm. This means that tasks like ad variant generation, product videos, and animated posters can now be handled more efficiently and creatively. Industries such as advertising, e-commerce, and gaming stand to benefit significantly from this innovation, as it simplifies the process of creating high-quality video content. The model's ability to understand and generate content from a unified context marks a significant advancement in multimodal AI capabilities.
## Feature Story
Supabase has launched Supabase Evals, an open-source benchmark that evaluates AI coding agents like Claude Code, Codex, and OpenCode on real Supabase tasks. This new tool is designed to test how well these agents can perform tasks such as building a schema, debugging a failed Edge Function, or fixing a broken RLS policy. The benchmark not only powers a public leaderboard but also supports an internal regression suite monitored daily. Supabase Evals is deployable today under the Apache-2.0 license and can be run locally using pnpm. It is particularly relevant for industries like developer tooling, cloud infrastructure, and regulated backends in sectors such as fintech and healthcare, where security is paramount. The framework evaluates agents across three dimensions: products, topics, and tasks. This comprehensive approach allows developers to assess the capabilities of AI agents in a real-world context, providing valuable insights into their performance and reliability. One of the key applications of Supabase Evals is in regression-testing documentation and skill edits, as well as gating SDK releases. By comparing agent harnesses head-to-head, developers can make informed decisions about which AI tools to integrate into their workflows. However, there are some constraints to consider. Local-stack runs require a Docker daemon, provider API keys, and specific ports to be free. Despite these requirements, the ability to run these evaluations locally offers significant flexibility and control to developers. Supabase Evals represents a shift towards more practical and applicable benchmarks in the AI coding space. Traditional benchmarks like SWE-bench have been criticized for not testing the right or valuable things, often being baked into the training data. Supabase Evals addresses these concerns by focusing on real tasks that developers encounter when using Supabase. This development is part of a broader trend towards more specialized and context-aware AI tools. As AI continues to evolve, the need for benchmarks that accurately reflect real-world applications becomes increasingly important. Supabase Evals is a step in this direction, providing a robust framework for evaluating AI coding agents in a meaningful way. Looking ahead, the impact of Supabase Evals could extend beyond Supabase itself. As more developers adopt this benchmark, it could influence the development of AI coding agents and the standards by which they are evaluated. This could lead to improvements in the accuracy and reliability of AI-generated code, ultimately benefiting developers and end-users alike. In conclusion, Supabase Evals offers a new way to assess AI coding agents, focusing on real tasks and practical applications. By providing a public leaderboard and an internal regression suite, it offers transparency and accountability in the evaluation process. As AI continues to play a larger role in software development, tools like Supabase Evals will be crucial in ensuring that these technologies are both effective and reliable.
What is Impact Vector: AI Tools?
Daily news about AI tools.