Skip to content

Methodology Changelog

This changelog tracks meaningful changes to how maestria-powered agents behave - what they do differently, what they’re stricter about, and what you can expect from them.

  • Comments and tests earn their keep. Implementers and reviewers now reassess explanatory comments and redundant tests before accepting a change.
  • Directive text consolidated. The human-facing output contract lives once in the shared global rules, and dense rules prose was split into scannable paragraphs; no obligation was moved, retired, or weakened.
  • Testing guidance is behavior-first. Agents reject tautological and change-detector tests, select test behavior before implementing, prefer end-to-end artifacts for complex features, and record failure-mode inventories before isolated-system implementation.
  • Reviewable increments and visible design. Agents commit coherent verified slices during multi-phase implementation, plan independently reviewable PRs early, and show architect or planner designs in a concise user-facing brief before dependent implementation. Integrated checks and review still govern delivery.
  • Project customization loads on every turn. Project-root .maestria/workflow.md then .maestria/rules.md apply as subordinate guidance across integrations; absent files leave defaults unchanged, and project content never waives safety or authorization.
  • PR and docs workflows are standalone skills. The create-pull-request and docs-update skills carry their procedures, while the core keeps a short obligation plus a pointer; see the CLI changelog for install and selection behavior.
  • Delegation terminology clarified as contract-driven delegation. Handoffs stay spec-aware and format-agnostic, with optional persistent intent refs via the spec-contract skill; docs clarification only, no behavior change.
  • Effort and process scale with the task. Investigation, verification, skill loading, and output structure track the size of the work; agents keep full assignment ownership and may add necessary in-scope regression coverage.
  • Skill references resolve at load time. Planner guidance no longer names the removed to-issues and to-prd skills, the duplicated Kimi orchestrator skill prescription is gone, and the pipeline states that Diagnose analyzes the bug, applies the minimal fix, and verifies the repair.
  • Implementation follows established engineering pressure. Builders validate and normalize inputs at trust boundaries, keep seams feature-local until repetition or shared change pressure justifies broader scope, trace every consumer of shared interfaces, and use executable sources or checks when several consumers must agree.
  • Refactors and migrations stay reviewable. Planners isolate enabling refactors with acceptance evidence and rollback points, prove migrations spanning many call sites or modules on a representative slice, migrate in verifiable batches, and give compatibility shims a removal condition or an explicit reason to remain.
  • Knowledge stays useful. Diagnose preserves only findings with durable future value, while Writer verifies claims against current code/config and gives operator-critical instructions runnable checks with expected signals.
  • Visual evidence on PRs where supported. For visual or behavioral changes, agents now attach a screenshot or short video after confirming the project targets GitHub (authenticated gh with --attach support) and a capture tool is available. Vision is optional verification, not a precondition.
  • Self-explanatory code over comments. Agent projections favor clear naming, small functions, and simple control flow; comments are reserved for concise context the code cannot express.
  • Shared human-facing output contract. Responses, status updates, briefs, comments, commits, PR metadata, and documentation never emit a Unicode em dash, while preserving code syntax, intentional literals, and quoted text.
  • Directive simplification (~24% leaner across canonical directives; ~21% as shipped in generated plugins). Canonical directives consolidated: the bounded-repair contract lives once in the global rules, specialist skill catalogs list only verified skills, and dead sync-config replacement anchors were removed across platforms.
  • Delivery is terminal-artifact driven. Delegated implementation outcomes complete at a pushed feature branch with an open PR; routine delivery requires no approval asks, while merge, release, and production actions remain separate authorization boundaries.
  • Transport failures are not verdicts. Failed or cancelled delegations get one adjusted-brief retry before structured blocker reporting; user-initiated or intentional platform cancellation stays terminal.
  • Full decision record: ADR-CORE-019 (docs/adr/core/ADR-CORE-019-directive-simplification.md).
  • Sessions continue through bounded repair. Normal engineering stays autonomous through continuation, scope-frozen repair, and reviewable PR delivery. Review and repair are bounded to material blockers, and incomplete specialist work becomes a structured blocker instead of an implicit user checkpoint.
  • Claude Code and Codex CLI projections documented. The Claude Code plugin is a native candidate, while the Codex CLI package remains provisional and skills-only. Both are installable through the maestria CLI’s host-native marketplace adapters.
  • Effort calibrated to task risk. Directives now match effort to stakes, prefer mature ecosystem solutions over custom infrastructure, converge reviews on material blockers, deliver routine engineering work through feature branches and PRs, and clean up task-owned background processes before completion.
  • Runtime-aware routing. Direct, focused, and full routes are selected according to task needs and host capabilities. fein can request the full pipeline, while direct execution remains available for known low-risk work where the host permits it.
  • Compact, evidence-oriented work. Directives and handoffs emphasize the material context, evidence, and outcome needed for the next decision instead of a fixed output schema. No-action and no-PR outcomes are valid when the work does not justify implementation.
  • Bounded progress and independent review. Repair loops use a practical normal bound of about three rounds and extend only with observable progress. Repeated causes require a strategy change or escalation; maker/checker review is used where the selected route supports it.
  • Clear lifecycle boundaries. Completion includes cleanup of delegated work, and nested agents do not duplicate an outer supervisor’s repository selection, scheduling, retry, or lifecycle responsibilities. Platform adapters may impose stricter execution limits.
  • Workflow contracts clarified. Routine validated commits on recognized feature branches remain autonomous after required review; push and later lifecycle actions stay separately gated. Bounded-repair and platform-enforcement notes were added, generated projections refreshed, and explicit mode reset behavior and read-only Sonar profiles retained where supported.
  • Directives re-aligned around evidence. Outcome, evidence, runtime authority, blind review, bounded repair, and autonomous routine work now share one contract. The host runtime decides whether work is performed directly or delegated, and nested supervisors do not duplicate scheduling or lifecycle work.
  • Directives streamlined. Universal contracts are centralized, and the directives add bounded autonomy, work-unit budgets, scope control, process lifecycle evidence, and checkpoint action boundaries.
  • Delegation is more selective. Handoffs stay concise, and Sonar research starts with the owning specialist, adding another specialist only when a distinct required output remains.
  • Selective routing by task class. maestria now guides turns through direct, focused, or full routes: explanations and discovery use direct execution, tiny edits use direct execution or a native builder, ordinary code changes use focused delegation, and complex or high-risk work uses the full pipeline with independent review where supported. The full route is selected for that work or for explicit fein requests, not forced for every turn.
  • Explicit modes and route-scaled safeguards. fein selects the full route, sonar is research-only, and blitz is a low-risk/direct bypass that does not waive safety floors. Delegation and review scale with the route. Maker/checker enforcement varies by platform; where separate sessions or tool controls are unavailable, the split is advisory.
  • Rule-override conditions and communication conventions. The orchestrator gained six explicit override conditions (user skip request, safety, mode override, frustration escalation, rule conflicts, explanation requests) and two shared rules: report errors matter-of-factly and lead with the action. Review triage chains directly into the commit flow after approval.
  • Blind review and fail-loud iteration exit. The reviewer no longer receives the builder’s handoff notes or self-assessment; it evaluates only the diff, requirements, and acceptance criteria. When a review loop runs three cycles with unresolved issues, the pipeline stops and escalates with a structured report that requires an explicit override to continue.
  • Directive recomposition. Core prompts were restructured with clearer sections and emphasis on critical rules, specialists gained structured handoff verification checklists, completion checks were standardized, a parallelization table and multi-lens review swarm were added, and prompt instructions became platform-agnostic.
  • Mode commands behave consistently across platforms. The fein/sonar/blitz mode commands now work the same way on OpenCode, Kimi Code, Pi, Cursor, and Hermes.

  • Compact agent directives. Agent directives compressed to cut orchestrator context usage by ~38% without removing rules or intent. Commit, review, and routing instructions are now grouped together. Rules shared by all specialists now live in one place instead of being repeated seven times. Fixes duplicate skill recommendations in diagnose and architect, a formatting bug, and stale cross-references.

  • Structured work results are now mandatory. After every task, the agent must produce a ## Changes summary table instead of a free-form description. This overrides the usual “write for humans” guidance for this specific output.
  • Documentation audit before every commit. The agent now checks four areas before committing: internal docs, the user-facing docs site, the user-facing changelog, and changesets.
  • Work results format simplified. The old 3-part narrative format (Overview, File-by-file, Cohesion) is replaced by a table-first format. Context paragraphs are now optional.
  • Change-type prefixes for easier scanning. Every entry now uses + (added), ~ (modified), or - (deleted). Breaking changes are marked with ! (e.g. !~). Test files are annotated with (test).
  • Work results appear in PR descriptions. The ## Changes section of every PR now includes the structured work results table, so human reviewers see the same summary without anyone having to write it twice.
  • PR descriptions follow a fixed structure. Every PR body uses Summary, Changes, Testing, and Breaking Changes sections.
  • Landing on main starts a feature branch. The agent no longer asks before checking out a feature branch when it lands on main; it creates one and continues.
  • Structured assumption tagging. Adventurer and architect outputs now require [verified] (confirmed from source) or [inferred] (best guess from context) tags on every assumption. Downstream agents can distinguish confirmed facts from inferred guesses.
  • Deterministic agents preferred. A new CRITICAL RULE mandates defined output contracts (checkpoints, success criteria, termination conditions) before delegating. Open-ended exploration must be scoped with time and resource limits.
  • Cognitive hygiene for delegation. The orchestrator now has a 5-trap pre-flight checklist (vague, overcomplication, attachment, rumination, overwhelm) to catch weak delegation prompts before dispatch.
  • Outcome specs over activity specs. Delegation now prefers specifying what to achieve over how to achieve it. Activity specs (step-by-step instructions) constrain specialist judgment; outcome specs let specialists apply their full capability.
  • Experiment framing for uncertainty. High-uncertainty tasks (unknown dependency, unvalidated approach, first exploration) can be framed as experiments with explicit hypotheses and termination conditions. The output is a validated finding, not shipped code.
  • Parallel speculation pattern. The orchestrator can now dispatch the same open question to multiple specialists with different lenses, then synthesize results before committing to a direction.
  • Unprompted incidental findings. Agents now surface relevant discoveries outside their brief after completing the primary task. Security, data, or production risks are flagged immediately.
  • First-principles decomposition. When blocked, agents decompose problems into verifiable sub-problems instead of trying harder with the same approach. Escalation path provided for sub-problems that resist decomposition.
  • Subagent permission alignment. All subagents received expanded bash allow-lists matching their actual tooling needs - read-only file operations pre-allowed, agent-specific tools added per role. Reduces unnecessary permission prompts from 70-88% of commands down to only unusual operations. Catch-all *: ask remains for safety.
  • Multi-lens review swarm. The orchestrator can now dispatch parallel reviewers with different focus areas (Security, Architecture, Performance, UX, General) for non-trivial changes. Each lens has exclusive scope and the orchestrator triages combined results.
  • Observation over reasoning. The reviewer now prioritizes running code and observing behavior over reasoning about correctness. Each review should include a command to produce visible proof.
  • Review triage pipeline. Issues are now categorized [fix]/[dismiss]/[escalate] by the reviewer, then validated by the orchestrator with conservative conflict resolution (when lenses disagree, the stricter categorization wins).
  • Docs audit before commits. Agents now check what docs need updating before every commit - no more waiting to be reminded.
  • Eliminate questions - autonomous philosophy. Agents now exhaust data, document assumptions, and proceed instead of asking for permission on every decision. Mid-phase questions (design choices, permissions, preferences) are eliminated across all specialists. Boundary checkpoints kept: commit autonomous, push conditional (auto on feature branches, ask on main), PR creation automatic on feature branches after push (user can edit after creation).
  • Autonomous commit protocol. Commit is fully autonomous - agent reads git log for past corrections, composes correct conventional commit message, commits without asking. Per-turn authorization requirement removed. Push is automatic on feature branches, asks only on main/master.
  • OpenCode permission alignment. Builder shell permissions expanded from 5 narrow patterns to a broad allow-list covering read-only file operations, git, package managers, and build/test tools. Diagnose edit permission changed from ask to allow. Both keep catch-all ask for unusual commands.
  • Structured work result summaries. Orchestrator now presents completed work as a file-by-file table focusing on signatures and interfaces (not function bodies). Builder reports at the signature level. Users can spot issues in 1-2 prompts instead of reading the full diff.
  • !!! Convention formalized. Non-negotiable rules are scar tissue from past failures, not preferences. New “Never delete what you didn’t create” rule. Recognizing User Frustration escalation ladder (5 levels with prescribed responses). Anti-Patterns enriched with named patterns and concrete fixes. Complexity-Based Routing (SIMPLE/COMPLEX classification). Session Flow section added - orchestrator proactively proposes next steps instead of going silent. Proven directive combos documented in COMPOSITION.md.
  • Writing style improvements. The “Write for humans” rule hardened to non-negotiable (!!!). humanizer skill added as always-load for orchestrator. Reviewer gained Writing Style checklist item 9 to catch violations.
  • Code intelligence tool preference. Agents now prefer code intelligence tools (MCP, CLI, or other) before falling back to grep and read when exploring your codebase. If you use a tool like CodeGraph, agents will route through it automatically.
  • Two new design principles. Agents now check from first principles before adopting existing patterns, and check for existing solutions before building from scratch. This pushes toward better design decisions with less wasted effort.
  • Anthropomorphism guard added. Agents no longer bias decisions based on “that would be a lot of work” - they operate at machine scale. The right technical choice wins, regardless of perceived effort.
  • Project-level workflow customization. You can now add .maestria/workflow.md and .maestria/rules.md to your project root to define custom delegation sequences and project-specific rules. Core rules always take precedence.
  • Single source of truth for agent directives. Agent prompts and global rules now live in one canonical location and sync automatically to all platform plugins (OpenCode, Kimi Code, Pi). No more drift between platforms.
  • Pipeline redesigned. Roles and their order now adapt to task needs instead of running a fixed sequence. The pipeline stops as soon as the verifier accepts - no unnecessary refinement loops.
  • Strict commit protocol. Agents now follow inspect → propose → execute → report for every commit. Commit and push authorization are separate - no autonomous pushes without explicit approval.
  • Workflow modes introduced. Prefix your message with fein, sonar, or blitz to control pipeline depth per turn:
    • fein - full pipeline, no shortcuts
    • sonar - reconnaissance only, no implementation
    • blitz - straight to implementation, skip preamble
  • Permission model overhaul. Fixed a bug where some agents couldn’t read the files they needed despite having permissions configured. Full audit closed gaps across all specialists. Delegation scoping strictly enforced.
  • Skill system. Agents now load skills based on what the current task needs, not a static preload. Four buckets: always load, load on trigger, defer to specialist, skip if. Skills stay out of context until they’re relevant.
  • Tool access hierarchy. Established clear rules for how agents find information: local tools first, then webfetch for known URLs, websearch only for discovery (with explicit justification).
  • Initial release. 7 specialist agents (orchestrator + adventurer + architect + builder + diagnose + planner + reviewer + writer) with global rules injection. Key patterns from day one: makers don’t review their own work, all delegation goes through the orchestrator, structured handoff contracts between agents.