| 01Computer use | Can Astra complete a browser or desktop workflow end to end? | Record the full task, every app handoff, required permission, human intervention, recovery step, and whether the requested outcome was actually completed. | Launch film + hands-on test |
|---|
| 02Coding | Does Astra improve real repository work, not only a staged demo? | Start from a concrete issue, then preserve the plan, code diff, terminal output, tests, regressions, review corrections, elapsed time, and final working state. | Launch claim + real repo |
|---|
| 03Agent orchestration | Does running several tasks together create leverage or just more supervision? | Show task boundaries, delegation, sequencing, shared context, conflicts, agent handoffs, completion status, and the moments where a person has to step back in. | Both official films |
|---|
| 043D & interactive worlds | How far can one prompt travel from an idea to a usable environment? | Compare the initial prompt with geometry, interaction, camera behavior, visual consistency, iteration count, export state, and the parts that still need manual craft. | Developer film + creator tests |
|---|
| 05Creative direction | Is Astra following taste, proposing useful alternatives, or averaging toward polish? | Keep the brief, references, first result, contrasting directions, revision language, rejected options, and the exact choice that made an output feel intentional. | Developer film |
|---|
| 06Benchmark audit | Do the launch numbers survive a reproducible task? | State the benchmark definition, inputs, scoring rule, model settings, number of runs, failures, variance, and whether your test measures the same capability as the headline. | Official claim + reproduction |
|---|
| 07Model comparison | Where does Astra win or lose against the strongest alternative? | Use the same prompt, tools, reasoning level, time limit, and success criteria; compare quality, latency, cost, consistency, and editing effort side by side. | Controlled side-by-side test |
|---|
| 08Reliability & recovery | What happens after the first wrong turn? | Leave failed attempts, repeated loops, self-corrections, retries, partial completions, lost context, and successful recovery visible before giving a reliability verdict. | Developer film + full capture |
|---|
| 09Access, speed & cost | Who can use Astra, how long does useful work take, and what does it cost? | Separate announced availability from actual access, then log wait time, generation time, token or usage cost, rate limits, and the value of the finished result. | Release details + first use |
|---|
| 10The AGI claim | Which demonstrated abilities support—or fail to support—the AGI label? | Define the term before judging it, map each criterion to visible evidence, identify missing capabilities, and keep capability, autonomy, reliability, and impact as separate questions. | Both films + explicit criteria |
|---|