To me, orchestration is not really the core problem. The bigger challenge is handling agent coherency over long periods of time, (drift) so my solution is a living forum-like system where agents shout to one another, document work in items, and follow doctrine handed down by the developer and work against a known specification that's checkable. I have used this workflow in several large projects that MUST be correct, and it works for me. Currently cleaning it up for other users.
Not saying this counts specifically but I love seeing devs who hate Jira build something like Jira to solve managing a team of agents. It’s a beautifully ironic thing to watch.
And I think most people who hate Jira actually just hate the process their company has imposed inside of Jira.
The vanilla out of the box product is actually quite close to old Trello, in simplicity. But there are so many ways to fuck it up and overcomplicate it, and so many companies proceed to fuck it up and overcomplicate it.
I have learned recently that I don't inherently hate JIRA, I hate our JIRA and byzantine process flows.
I think because other task trackers like Trello are simpler and have less config options (per my last use, long ago, pre-Atlassian-acquisition), it's impossible knock on wood to get into the antipatterns I've seen in those tools.
Ironically I think one of the main goals of management I have seen consistently over time - ensuring tickets are accurate, up to date, and reflect reality - would end up much more true if the task tracking tools were just much simpler, with fewer fancy features.
It might be you'd like Epiq. It is an issue tracker I built that actively does what it can to not be another Jira. It has way less fields and configurations per ticket (description, tags, assignees, attachments, comments), and it integrates with your code/repo, it is local-first and distributed so super fast.
https://ljtn.github.io/epiq/
Yeah, my company has a simple, straight-forward system, but I still dread using Jira. And I love Linear, so it's not general ticket management system hate! Jira just sucks.
>> And I think most people who hate Jira actually just hate the process their company has imposed inside of Jira.
I think what you describe is a problem for a lot of people, but what isn't acknowledged is that some people don't want any kind of process, or rather, they don't want to deal with any kind of friction. These people don't like Jira without the customizations. They don't like equivalent tools.
Given these tools regiment the world to some extent, many programmers are fine with the minimal friction. However, you have big problems when people who don't value any process get into positions of authority. This is also often a problem when you interface with other units of an organization who don't understand the value of a process.
I hate Jira because it’s one of the slowest applications I’ve used since the i486 days and the UX is designed to force you to click across 3-4 different fields and modals to accomplish an action that should be one form.
It’s business software for business people that need to be guided through SOP because they have so many different workflows, not engineers that have streamlined work processes to reduce mental load on non-engineering work.
Would you be more inclined to use Epiq? It is instant, as every write is local, synced over Git. It is built with engineers in mind. I'm its author for disclosure.
That actually looks pretty incredible, including some features I never knew I wanted.
Unfortunately I work for a very large corporation and don’t have the pull to get something like Epiq brought on - but I am definitely going to evaluate it for my non-FTE work.
As a side note - it was impossible to find your project through search engines. I only found it by searching on github itself.
Hey, thanks for the kind words. I need to get to the bottom of Epiq not showing up in search engines! It doesn't help that Skoda released a car model with the same name a month or so after :D. Totally understand the large corp dynamics!
Yup, I literally just hate their web UI. Once I found the Atlassian-blessed jira "cli" [0] (after searching for a long time for a tui, which is actually what the "cli" is primarily) I have actually be fine with using Jira again. As an added bonus, being a cli has made it easy to interface with using my AI tooling.
2nd this but I did need to write a session start hook to prune out most of the skills so it was just the non Rovo ones. Took about 15min to setup though so I am not complaining. twg works really well.
I think we'll see it more and more... open-ended Q&A / novel work is everyone's first experience of LLMs, but I don't think it's actually representative of most of the work people do. Coding kind of sits in the middle as it's novel but semi-structured.
Chat is great for novel open-ended work, but is terrible at concurrency - i.e. managing > 1 agent runs. As agents get cheaper, I suspect we'll see more more bounded processes where you want to run lots of items through them at once, and thus Jira-like things... (we ended up having this problem and building something like this, albeit more medieval themed [0])
We also found the other piece that Jira-style solves (and that chat makes super hard) is multi-player. I've yet to see a good implementation of a multi-player chat solution to agents working concurrently; and we found that an async model is the only sane solution to this.
I have been using Backlog.md in precisely this way and it's very effective. The finished tasks end up being a knowledge store that the git history doesn't handle. I just finished a really large project that was planned, specc'd out, and broken up into backlog tasks. As I went through each task I stored all the changes and business logic issues in the task. Having a time stamped record of the thought process and experiment results has been very useful. https://github.com/MrLesk/Backlog.md/
Haha. Locally it was passing. That badge is GitHub running the tests on its own machines, where a couple of new browser tests fail. Fix is going through the same loop: one agent fixes, another verifies, then it merges.
I'm trying to fix the "agent said work was done, it was not" problem by defining work as runnable specifications in the form of a pytest suite.
Work is first defined in new or updated specifications, then the change gets made. You have to resist making natural language Todo lists and have the agent write runnable unit tests. My project is a CLI tool so it's fairly easy.
Verifying if it's done is a matter of running the specifications test suite.
In a totally not original way, I'm testing this workflow in my own agent orchestration tool. I know, there's just so many already. I'm building it for myself.
So far, the system is holding up but it's way too early to declare it a success.
I think the same idea could be applied to other projects in different domains.
With this workflow I can read and update the specs, and run it against a built binary and be much more confident that the code works. Claude code has picked up the system without complaining for the most part.
We do have what you call Items and Shouts in the form of an agent communication system and a Linear style issue tracker that agents and humans use together.
In our case everything around orchestration is discussed with the agent itself and they take action for the user via our CLI. We found any UI to prescribe how teams of agents should coordinate too clunky. So in our case you just tell the agent: "every codex luna max implementer gets an opencode muse spark reviewer. when a subtree is done, review with opus high.". It then sets up the graph and the system enforces the rules.
How do you handle the context? As in: you have a forum like system as you say, how are agents finding the correct "posts"? How do you handle drift or specs that change over time?
Code means a bit of mental engagement, more than yaml, json etc. You either review that or not review that, but it's an addition. While I am not touching anything made by docker-the-bloater anytime soon, I could think that at least.
Go doesn't have a good mod/plugin story for harness devs to provide to their users
I say this as a gopher who also has a custom harness written in Go, it's a real challenge compared to the TS setups, but TS plugins are also a security concern
I agree! Comes with the territory of a compiled language I suppose?
I haven't given this much thought (I also just started with golang last year) but I'm assuming the only real path here is to provide extension SDKs in Lua or other languages whose interpreters have been implemented in go.
I think that's what grafana's k6 [1] does with a JS interpreter
I was thinking I recalled something like that, but the nice thing about the TS ecosystem is that you can write the same language, importing and using the harness sdk types.
Go has the plugin package, but I haven't actually seen anyone use it (outside of toys and blog posts)
My custom harness has a dagger based environment, binary runs in the host, every tool call and file modification creates a new layer in a container - for better session time travel and forking
It may work for some things, but you really want to be wired into that system to do the more interesting things.
I didn't say it couldn't work, it's just way more complexity and effort with IPC or Go than TS. That's one of the reasons harnesses are choosing TS. It's also a very popular language.
So glad I never had to deal with Angular, especially v1. I tried learning it and it just seemed like it would actually make applications more incomprehensible than before. React originally addressed the actual problem that most frameworks mistook for not being that important. I still don't particularly like React, but it makes sense why it became popular.
Cursor is easily the worst experience I've had with an AI harness, yet it was acquired for billions of dollars in spite of being a middleman tied to a ripoff of VS Code. None of the colleagues of mine who were touting it last year are still using it. But if you can make yourself look like the next big thing, you'll get money thrown at you. Just look at Omarchy.
A good experience necessarily takes time and careful thought. Nobody these days can do that while also remaining relevant, or even surviving.
> Cursor is easily the worst experience I've had with an AI harness
I kinda like cursor, mostly because it provides a nice review UI and I can use whatever model I want.
Getting away from all the Claude-speak has been a revelation, and fortunately Anthropic helped me here by hyping up GLM 5.3. Like, if there's an open weight Fable competitor, then why wouldn't I use it?
What do you prefer to Cursor and why? Just asking because I don't really experiment w/ AI harnesses much outside of work, where cursor has a lot of traction, and I'm too burned out to explore much these days
I get your point but cursor is fine. And it might be an unpopular opinion but grok 4.7 is also good inside cursor. So yes it might be acquired for a lot but it kinda worked for both parties well.
With the advent of LLMs my one true hope for humanity was the final death of the "JS/ilk things" like Electron, React Native, other things like that, and the culture of "everything is TS/JS or can be turned into that". Looks like it's rather growing exponentially.
Docker Agent is a harness. There is a sandbox mode that can be used to run it in docker sandbox (a VM, not a container). If you don't use sandbox mode then I assume it is running in a container.
If you don't want to use their harness then you wouldn't use docker agent and instead use their `sbx` cli to run the harness of your choice (claude, codex, pi, etc).
I think it will be great if Docker can get people used to using secure VMs. I am developing a similar project (still a work in progress): https://github.com/gregwebs/agent-vm
I've found sbx to be very helpful. really like the sentinel value wrapper they have going so you can add secrets but the model can't see them. Outbound calls get looked up by sentinel value and the real one goes out to whatever API you're auth'ing too. There's other features but just having claude run in a sandbox and easily see what it does and doesn't have access to has been great.
I just wish sharing whitelisted domains between sandboxes were easier or more intuitive. Last time I tried doing this with sbx I was handling their yaml files by hand.
We built something similar, the ideas are strangely converging for production it feels like. But we used routing rather than direction agent invocation of agents. It gets exhausting to keep track of large number of concurrent agents. Best to keep a single orchestration agent that can route to agents based on capabilities.
Docker Agent could be beneficial for research studies. For example, while writing a ML paper about agents I need to rerun experiments so that:
- others can repeat my results
- validate that variability from LLMs is not causing overconfidence in the results
But currently uv is good enough to create isolated, repeatable environments. And it isn't limited to Linux like standard Docker. So when I want to test small samples on my local PC's GPU (Windows or mac) before running on beefier hardware, I can do that easily.
I mean, if you're going to do that at scale, not that you are, but something like https://github.com/kvcache-ai/AgentENV is much more suitable for research purposes
I tried really hard to make this work for me a couple of weeks ago but it was the brittlest harness out of any that I've used. I'd come back when it's more mature.
Sessions just broke all the time for me. The one time I dug into it, it ended up being a known issue where if codex returned over 10k characters for a turn it breaks the session. Because docker agent does not use protocols like ACP, instead trying to parse Codex's internal state in a brittle manner.
I tried not to do anything too crazy, but even use using an API key in the docker agent natively was getting me sessions that would hang and couldn't be recovered.
Not the most common use case around them: likely meaning that the Codex harness isn’t used that often internally at Docker, is what I would surmise, at least that is how I read it.
I work at another big tech company, and up until recently that would’ve been true for us as well (but it’s changed now, we partnered with OpenAI)
Last year (2025), I lost count of how many times I tried using an LLM to set up Docker. It was still the era of prompting, then copy-pasting the response and seeing if it worked... and it almost never did. It’s great to see this new Docker agent. For me, it would be even more appealing to see an agent and a model specialized in Docker working together, capable of proposing advanced configurations and getting them working on the first try.
Modern models should be better with mainstream technologies like docker engine or docker compose, but still to try to help them we started to work on distributing skills ( https://github.com/docker/skills ) in addition of our specialized agent (gordon) integrated in various products.
I literally built this for self hosting apps at Tarvis.io
You can drop any GitHub repo and the agent builds and deploys the app to your Tarvis workspace along with taking care of ssl domain and any other configuration and making sure the app is healthy.
142 comments
I've been working on Pullboard and just open sourced it: https://github.com/pullboard-dev/pullboard
To me, orchestration is not really the core problem. The bigger challenge is handling agent coherency over long periods of time, (drift) so my solution is a living forum-like system where agents shout to one another, document work in items, and follow doctrine handed down by the developer and work against a known specification that's checkable. I have used this workflow in several large projects that MUST be correct, and it works for me. Currently cleaning it up for other users.
The vanilla out of the box product is actually quite close to old Trello, in simplicity. But there are so many ways to fuck it up and overcomplicate it, and so many companies proceed to fuck it up and overcomplicate it.
I have learned recently that I don't inherently hate JIRA, I hate our JIRA and byzantine process flows.
I think because other task trackers like Trello are simpler and have less config options (per my last use, long ago, pre-Atlassian-acquisition), it's impossible knock on wood to get into the antipatterns I've seen in those tools.
Ironically I think one of the main goals of management I have seen consistently over time - ensuring tickets are accurate, up to date, and reflect reality - would end up much more true if the task tracking tools were just much simpler, with fewer fancy features.
So what they really hate is their company. Which is usually fair and understandable.
I think what you describe is a problem for a lot of people, but what isn't acknowledged is that some people don't want any kind of process, or rather, they don't want to deal with any kind of friction. These people don't like Jira without the customizations. They don't like equivalent tools.
Given these tools regiment the world to some extent, many programmers are fine with the minimal friction. However, you have big problems when people who don't value any process get into positions of authority. This is also often a problem when you interface with other units of an organization who don't understand the value of a process.
It’s business software for business people that need to be guided through SOP because they have so many different workflows, not engineers that have streamlined work processes to reduce mental load on non-engineering work.
Unfortunately I work for a very large corporation and don’t have the pull to get something like Epiq brought on - but I am definitely going to evaluate it for my non-FTE work.
As a side note - it was impossible to find your project through search engines. I only found it by searching on github itself.
[0] https://github.com/ankitpokhrel/jira-cli
Need a specific feature? Plug that in later when you actually need it.
Chat is great for novel open-ended work, but is terrible at concurrency - i.e. managing > 1 agent runs. As agents get cheaper, I suspect we'll see more more bounded processes where you want to run lots of items through them at once, and thus Jira-like things... (we ended up having this problem and building something like this, albeit more medieval themed [0])
We also found the other piece that Jira-style solves (and that chat makes super hard) is multi-player. I've yet to see a good implementation of a multi-player chat solution to agents working concurrently; and we found that an async model is the only sane solution to this.
[0] https://pumpup.com
https://yegge.ai/gastown
> gate: failing
Gave me a good chuckle. Cool idea though, and I like the records. Reminds me a bit of the OpenAI hacking incidents.
Thank you for coding something that I intended to code but have been too lazy to do.
Work is first defined in new or updated specifications, then the change gets made. You have to resist making natural language Todo lists and have the agent write runnable unit tests. My project is a CLI tool so it's fairly easy.
Verifying if it's done is a matter of running the specifications test suite.
In a totally not original way, I'm testing this workflow in my own agent orchestration tool. I know, there's just so many already. I'm building it for myself.
So far, the system is holding up but it's way too early to declare it a success.
You can check it out here:
https://github.com/egzo-ai/egzo/tree/main/specs
I think the same idea could be applied to other projects in different domains.
With this workflow I can read and update the specs, and run it against a built binary and be much more confident that the code works. Claude code has picked up the system without complaining for the most part.
We also open-sourced a similar system called https://github.com/madeinorbit/podium
We do have what you call Items and Shouts in the form of an agent communication system and a Linear style issue tracker that agents and humans use together.
In our case everything around orchestration is discussed with the agent itself and they take action for the user via our CLI. We found any UI to prescribe how teams of agents should coordinate too clunky. So in our case you just tell the agent: "every codex luna max implementer gets an opencode muse spark reviewer. when a subtree is done, review with opus high.". It then sets up the graph and the system enforces the rules.
It works really well for us.
How do you handle the context? As in: you have a forum like system as you say, how are agents finding the correct "posts"? How do you handle drift or specs that change over time?
How is "no code required" a selling point when we all have agents that can almost instantly puke out all the mostly-correct code you could want?
All these translation layers exist because the computers only understand 0s & 1s. Imagine talking to a person who only understood binary numbers.
All the cool kids have one!
I say this as a gopher who also has a custom harness written in Go, it's a real challenge compared to the TS setups, but TS plugins are also a security concern
I haven't given this much thought (I also just started with golang last year) but I'm assuming the only real path here is to provide extension SDKs in Lua or other languages whose interpreters have been implemented in go.
I think that's what grafana's k6 [1] does with a JS interpreter
[1] https://github.com/grafana/k6
Go has the plugin package, but I haven't actually seen anyone use it (outside of toys and blog posts)
https://oneuptime.com/blog/post/2026-01-25-plugin-system-go-...
why not just use an SDK where you are using a plugin?
It may work for some things, but you really want to be wired into that system to do the more interesting things.
jk its trash
Vue and svelte are easy to use and learn/adopt. But look what won. React.js.
So whatever the ai agent with the vast user basae will win regardless.
---
I know it's a slippery slope, but given you brought up JS framwork analogy, I had to fall into it
Now Angular v1, man… if I ever see another digest loop error in my life, I will scream
Cursor is easily the worst experience I've had with an AI harness, yet it was acquired for billions of dollars in spite of being a middleman tied to a ripoff of VS Code. None of the colleagues of mine who were touting it last year are still using it. But if you can make yourself look like the next big thing, you'll get money thrown at you. Just look at Omarchy.
A good experience necessarily takes time and careful thought. Nobody these days can do that while also remaining relevant, or even surviving.
I kinda like cursor, mostly because it provides a nice review UI and I can use whatever model I want.
Getting away from all the Claude-speak has been a revelation, and fortunately Anthropic helped me here by hyping up GLM 5.3. Like, if there's an open weight Fable competitor, then why wouldn't I use it?
React was cool before the hooks and "use whatever" all over the place.
"Uh, OK, so how do I get this thing working? Where are the docs?"
"Well, first you install its unique framework..."
Which will probably also be complex and a bit slow and a bit hard to learn because of all the big tech company problems it solves
If, like me, you couldn't find any security-related info on the linked page.
If you don't want to use their harness then you wouldn't use docker agent and instead use their `sbx` cli to run the harness of your choice (claude, codex, pi, etc).
I think it will be great if Docker can get people used to using secure VMs. I am developing a similar project (still a work in progress): https://github.com/gregwebs/agent-vm
But currently uv is good enough to create isolated, repeatable environments. And it isn't limited to Linux like standard Docker. So when I want to test small samples on my local PC's GPU (Windows or mac) before running on beefier hardware, I can do that easily.
I tried not to do anything too crazy, but even use using an API key in the docker agent natively was getting me sessions that would hang and couldn't be recovered.
I work at another big tech company, and up until recently that would’ve been true for us as well (but it’s changed now, we partnered with OpenAI)
Sorry, what do you mean there?
You can drop any GitHub repo and the agent builds and deploys the app to your Tarvis workspace along with taking care of ssl domain and any other configuration and making sure the app is healthy.