Found 38 results
ah nice, I've actually started making a new harness on top of jolt, but it's not quite ready for general use yet https://github.com/yogthos/samizdat/ But I wonder if you could integrate Clojurust with dirge now, it's already a pretty solid base, and you could swap out Janet with Clojure for plugins and extensions. That could be a neat project to try.
Oh I meant having the harness repair parens automatically before writing code, seems to work well for me in dirge.
Sweet! so for now, would it make sense to drive it from dirge?
I'm still tuning it right now, for example it doesn't have a TUI yet, but yeah my plan is to turn it into a replacement for dirge
@U050CBXUZ is samazdat at a place now where you use it instead of dirge? or do you use both for different things, perhaps?
Dirge actually works like a regular code harness where you can tell it to spawn subagents, and samizdat is an iteration on that I'm building now where I decided to try splitting things automatically. And the agent decides, it's asked to look at the task and decide whether it should be done as a unit or if it's better to split it up, and if it tries to solve a task and keeps failing it's nudged to split it up. Why should this be a unique thing that a human decides over any other thing you let the agent decide?
@U050CBXUZ ok, so the dirge is about splitting the work into smaller pieces tasks/subtasks. Who decides what is worth doing? Shouldn’t it be a human deciding?
@U050CBXUZ ok, so the dirge is about splitting the work to smaller pieces tasks/subtasks. Who decides what is worth doing? Shouldn’t it be a human deciding?
The problem I found in dirge is that folding over context only helps to an extent. You still end up with a lot of noise in form of implementation details, and you can curate the context as you go, but there are diminishing returns. Simply not putting things in the context that the agent doesn't need to worry about is a better approach in my experience.
I landed on this concept back with dirge, and I found it very much improved what the small models could do
Dunno if dirge, but https://github.com/Blockether/vis you can :)
Can I use dirge with a subscription from openai or anthropic?
Yeah, I've been doing mine in Rust and I'm feeling the same pain you did with dirge.
It's just issues on the repo for dirge, but I did start building a new harness in Jolt with a lot of the lessons learned from dirge https://github.com/yogthos/samizdat/ and the goal with it is to have a workflow that can be modified at runtime and adopt to the project being worked on. So, general features like talking to a provider, running tools on the machine, using the db, go in the src layer that gets built, but any behaviors such as the agentic loop itself go in workflow manifests that can be modified at runtime. The harness defines a supervisor role whose job is to observe the agentic loop running and then intervene to fix problems such as the implementer model getting stuck or not finishing work. My idea is to package some general workflow templates that will get copied per project and then evolve alongside each project organically.
Thanks for sharing. I've been trying out dirge lately and was wondering if there's a channel or similar for the project. Is it just the issues etc on the repo, or is there another space?
eventually I'll port other providers from dirge, but just hasn't been a priority since these are the two I use
that's how dirge works right now
and in dirge these were all just add hoc things I'd be tack on to the loop as it evolved organically
@U050CBXUZ is jolt using dirge or some harness? What kind of harness does one want to use for project maintenance, after most of the features have been done, for a project like a language? Public facing libraries seem to need different harness needs than more ephemeral applications or green field prototypes.
So, the progress I'm making on samizdat compared to dirge is mind blowing for me. On the one hand, it's always easier to do things the second time around when you know what you're doing a lot better. But at the same time, the REPL driven loop and the agent being able to audit itself and tweak errors as soon as they appear makes an incredible difference. Dirge ended up with a lot of cruft initially because the tests weren't rigorous enough to catch behaviors that were not wired up correctly, and exercising them from the UI was a pain, but when you have a REPL, ensuring that things really do work end to end is trivial.
nice, yeah I'll see how it might fit with dirge here
Have been playing with pi, claude and codex. Was going to use config.edn to add some definitions for dirge, but ran into a blocker -- some of the monitoring code relies on upstream herdr agent ... cli commands, requiring herdr to support a particular harness. Which may limit its relevance for you @U050CBXUZ
oh haha, I forget what one I used, I stopped bothering because I have it built into dirge now, and I often just farm out implementation to it anyways
I've basically been doing a/b testing with dirge when I implement features, where I get the model to implement a multi step task and see which way it deos better
One big part is having persistent knowledge across sessions. For example, dirge keeps sqlite db where it distills memories and keeps them up to date, this also makes it possible to do auto compactions, because useful information isn't lost in the process. In terms of plugin customizations, one common one I have is the repl workflow, where the harness puts the model into a repl loop automatically. Another one I use commonly is to set up the agentic loop around the steps I want the model to do for a specific project, which tests it needs to run, what steps to take when making releases for it, etc.
yeah I think he might be on to something there, this is kind of what I'm finding with dirge too, the big wins turned out to be project specific memory and the ability to write plugins tailored to the project using Janet
So dirge is creating halting programs in prose for it's sub-agents? That's nifty.
Re. mixing and matching harnesses - I've spent quite a lot of time/tokens over the last week or so on a harness-agnostic subagent skill using herdr that I'm enjoying. Had built up a workflow in pi that I liked, but made it hard to evaluate other harnesses (e.g. dirge) on real work without porting my workflow over. So refactored to use generic skills with babashka scripts in the skill tree to reduce tool call overhead for repetitive/mechanical stuff, rather than harness-specific extensions. Seems to work well so far - pi and codex happy spawning each other as subagents etc. Inter-agent comms contract defined in terms of herdr pane management/interop. Caveat - my use of subagents is serial or very low fanout - about context and cost management rather than managing swarms of agents.
@U11BV7MTK and new dirge is out if you want try see how it compares with opencode
and new dirge release is just building here
meaning claude code uses dirge as a sub process? or do you have fable write a markdown file or something and then spin up dirge separately?
that's what I've been doing a lot actually, I added headless mode to dirge and I get fable to do planning, and then have it use dirge as mcp to implement
might be something like fable on max effort/reasoning for a plan and then delegate for execution. so store a bunch of plans. And now start investigating harnesses. opencode was the default so i’m really stoked that you suggested dirge ( i keep wanting to say dirigeste from the tellman threadpool lib)
going to get to dirge in a bit
dirge integrates task tracking right into the harness using sqlite, so the model always has a list of issues, and then todos are active tasks it's working on, and the current task gets injected at the start of each round to keep it focused
actually let me get an update to dirge out that I'm working on right now, it should improve long horizon goal tracking
i’m happy to try this out. I’m looking for an established workflow and had just defaulted into open code. i’ll happily try out dirge. Ideally my plan takes about 13 cents to implement and then i can compare deepseev-v4-flash between dirge and opencode
haha I got claude to read the paper and see what might apply for dirge 🙂
Up front: the transfer is analogical. RAX is a symbolic planner + executive with a formal domain model; its specific techniques (Latin-squares generation over model parameters, model-based mode identification)
don't port. What ports is the control structure — the failure ladder, fidelity tiers, progress-vs-budget, residual replanning. That's where I focused.
What dirge already does well (paper analogues it has)
- Reflect-then-pivot on repeats — http://storm.rs suppression + REPEAT_LOOP_GUARD (http://run.rs:2108), with http://reflexion.rs accumulating abandoned approaches. This is EXEC's "attempt an alternate method."
- Recovery checkpoint on error streaks — http://failure_tracker.rs at 3 consecutive errors, with timeouts weighted double via http://activity.rsOutcome. Analogous to MIR's recovery request.
- Independent judge at the finalization boundary — http://critic.rs / http://goal.rs / http://code_review.rs with a strict precedence chain (FollowUpSource).
- Safety net — permission checker, deny_tools, plan-mode read-only lock, sandbox. This is the RAX manager, and the paper is explicit that this is what bought the project's confidence to fly at all.
- Rewind — http://snapshots.rs captures pre-mutation file content per user turn.
The loop is already better instrumented than I expected. The gaps are specific.
Findings
1. HIGH — No "abort to safe state" rung. EXEC's ladder is: alternate method → request recovery → cleanly abort the plan, bring the spacecraft to a safe state, request a new plan (§2). dirge's ladder stops at rung
2. Both existing guards keep the model digging in a possibly half-mutated tree; nothing says "this attempt failed, restore the last known-good state, re-plan from there." The machinery exists —
snapshots::restore_from — but is user-driven (/rewind) only.
2. HIGH — Every guard keys on errors; none on stalled progress. The dominant late PS failure was "operating correctly but unable to find a plan within the allocated time limit since its search was thrashing"
(§4.4). Not an error — a non-result within budget. storm needs identical calls, failure_tracker needs errored results, context_depth needs repeated same-file touches. A model making successful, varied, useless
calls trips nothing until max_turns hard-stops at http://run.rs:2529.
3. HIGH — Verification is one flat binary gate, not a fidelity pyramid. §4.2 is the paper's most transferable engineering result: cheap tier gets volume (200 planner variations, overnight, 7:1 speed), expensive
tier gets only nominal + timing — licensed by the interfaces being identical across tiers, only response fidelity differing. Plus §4: developers front-line testing during integration found most bugs; formal
high-fidelity testing "found few." dirge's http://verifier.rs asks one question — did a build/test command run and pass — and only at finalization. cargo check and cargo test --all-features are indistinguishable to it.
So verification is end-loaded and untiered.
4. MED — The todo list is planned once, never re-planned against observed state. MM plans one horizon at a time folding in EXEC's projected state, and the second plan contains an activity to generate the third
(§2, §3.2). write_todo_list is bulk up-front; afterwards items are only checked off under a "finish your todos" nudge. Nothing folds what was learned in items 1–3 back into items 4–10.
5. MED — No goal triage. Rejecting low-priority unachievable goals is a validation objective for PS, not a failure (§3.1b) — optical-nav windows had time for only a subset of asteroids. dirge's todo items carry no
priority, and the nudge pushes completion of everything; cancelled is supported by the store but not invited by guidance.
6. MED — Early termination produces no residual-objective handoff. The 2-day scenario aborted at 70% of objectives; within 10 hours they built a 6-hour scenario targeting precisely the remaining 30% and reached
100% (§3.4, §5). dirge emits a truncation notice; nothing states which acceptance criteria were met vs outstanding, so a resumed session re-derives scope.
7. MED — Assumptions are implicit and never re-checked, and divergence isn't attributed to inputs. PS testing assumed turns ≤20 min; real turns exceeded an hour (§4.4). More pointedly: in flight, PS followed an
unexpected search trajectory because the goals file on the spacecraft differed from the testbed — it was solving a different problem (§5). debug.md goes straight to hypothesizing about logic; there's no "confirm
your inputs are what you think" step.
8. MED — "Check for edge cases" is generic; the paper's method isn't. Pairwise coverage missed a bug depending on an equation over three continuous parameters, because the generator was independent of the domain
model. The fix was boundary analysis derived from the model: 25 boundaries where plan topology changes → 88 cases → 2 bugs, plus a quantified residual (0.5% of start times) with a contingency procedure (§4.4).
9. LOW — No change-control ramp late in a run. The CCB rejected 6 of the last 10 changes: "every bug fix modifies a system that has already gone through several rounds of testing," and the alternative was
restricting the scenario so the bug is never exercised (§4.3). dirge is uniformly eager all run. Note this does not conflict with your "fix every bug you find" rule — RAX presented every change with its specific
lines for a decision. The dirge form is surfacing, not suppressing.
10. LOW — Budget is enforced but invisible to the model. RAX had explicit allocations (32 MB, 45% CPU) with measured peaks (29 MB) (§4.1). dirge knows turns-used and context%; the model doesn't. A silent hard stop
can't trigger triage; a visible countdown can.
11. LOW — Concurrency changes get the same evidence standard as everything else. A missing critical section deadlocked in flight after "thousands of previous races" passed on the ground (§5). Green tests are
currently sufficient evidence for any change.