babashka 2025-09-02

I'm deciding if I'll change the non-daemon thread behavior in bb. So far they were just killed (as an oversight on my part) because of a (System/exit (if ... 0 1)) somewhere at the end of babashka's main function. This is the reason you had to insert a @(promise) in cases where you e.g. started a webserver like:

$ bb -e '(org.httpkit.server/run-server {})'
#object[clojure.lang.AFunction$1 0x64996254 "clojure.lang.AFunction$1@64996254"]
but in "normal" Clojure this is unnecessary. My proposal is to switch to the behavior that clojure -X has: it switches to an executor with daemon threads for futures and agents, so spawning a future doesn't make your script wait at the end for the executor to stop or by calling (shutdown-agents). This is the diff: https://github.com/babashka/babashka/pull/1858/files • ❓ Would this break any scripts and in which cases? • ✅ Would this improve Clojure compatibility? As an intermediate solution I could put this behind a flag or env variable and then after some feedback I can switch this to the default behavior (or not).

🤔 1

@alexmiller in https://clojure.atlassian.net/browse/TDEPS-198 in the decision matrix shouldn't "On startup: make future/agent threadpool into non-daemon threads" be: "into demon threads" ?

And another question: in the third row: how is it that shutdown agents break agent/future behavior compared to making them run in daemon-threads? Or does broken mean: broken in the sense that the -X invocation doesn't quit when you have a long running future somewhere?

The reason I'm asking is, I might have to do a shutdown-agents at the end in addition to the daemon thead executor on startup since somehow the linux static binary doesn't quit when using an agent. I have yet to dive into this. Ostensibly the behavior would still be the same when doing both, right?

When I just add (.shutdown clojure.lang.Agent/pooledExecutor) in the linux static binary it starts working correctly. Maybe I can fix this by re-initializing the pooledExecutor myself or so. Note: the linux static binary has all of its logic running inside a new thread, other than the main thread. Perhaps that's somehow relevant.

Just to clarify, although Clojure has that behavior for -X, for invocations via -M -- including -e -- the executor is not switched. The behavior for -X was specifically changed so that it matched programmatic invocation of the underlying fn in a pipeline of calls, whereas -M invocation was assumed to exit and didn't need to worry about daemon threads, if I recall correctly.

Let me clarify. • The JVM waits for non-daemon threads. • clojure's future / agent etc uses a thread pool that creates non-daemon threads. when you use a future in clojure, you usually have to call shutdown-agents too (even when the future finished) which is annoying for scripts. clojure -X switches the behavior of future / agents to daemon-threads. • babashka never waited for non-daemon threads (due to an oversight of calling System/exit regardless at the end of main). I think the clojure -X behavior makes most sense for babashka, keeping the convenience for scripts with future not causing the script to take 30s (or so?) to quit, but still waiting for non-daemon threads (which happens on the JVM and with Clojure always)

this was the original issue for clojure -X: https://clojure.atlassian.net/browse/TDEPS-198

Alex's comment in the ask question was: Sean's analysis is mostly correct. Doesn't really have anything to do with chaining, which is not yet a feature, but just in wanting to match expectations on shutdown when executing a -X function. There are multiple cases here: * Use future/agent, have no background threads, expect full exit. Will wait for 1 minute for background agent pool to shutdown, so MUST have shutdown-agents or System/exit. * Have background threads (like socket server). Expect to block and NOT exit. CANNOT use System/exit (or you kill server). * Use future/agent, and have background threads (this case). Expect to block and NOT exit. CANNOT use System/exit (would kill server) or shutdown-agents (would kill futures/agents run from server). I don't think there is any behavior that works for all cases by default (you can always wrap a function and do whatever you know is right).

Right, which was why -X and -M have different behavior -- and there was discussion of chaining but it was elsewhere (and never went anywhere -- we got tools.build instead, essentially 🙂 )

Not a big deal. I just wanted to mention that it didn't happen on all entry points, only on -X.

I don't understand the "why" in "which is why -X and -M have different behavior", why exactly? My assumption was that -M doesn't have this behavior since it's old and cannot be changed, but -X is pretty new and therefore could apply this improvement

to me -X and -M -m don't differ that much conceptually, they both invoke a function, just the arg parsing is different

I'll just ask @alexmiller this specific question for clarification: Why does -X use non-daemon threads for future/agent and -M (-m) does not?

Alex Miller (Clojure team) 2025-09-02T18:31:25.918719Z

That would be a change in Clojure itself

Alex Miller (Clojure team) 2025-09-02T18:32:43.828849Z

Because clojure.main is in Clojure vs the - X stub that is injected by the CLI

Would you say the -X behavior is overall more desirable for scripting?

And if you could turn back the clock, would you like Clojure -M -m to behaved this way?

Alex Miller (Clojure team) 2025-09-02T18:34:54.913009Z

I’d prefer different behavior in Clojure itself

Alex Miller (Clojure team) 2025-09-02T18:35:47.080679Z

Which maybe will change some day :)

different behavior as in?

do you mean, move the -X daemon behavior in Clojure itself?

Alex Miller (Clojure team) 2025-09-02T18:39:19.893889Z

I’d rather that the default thread pools behaved differently. We are considering changes/additions to some of that for vthread support

Alex Miller (Clojure team) 2025-09-02T18:39:51.502879Z

Then you wouldn’t need anything special for -M or -X

so with those tweaks, -X still wouldn't need (System/exit 0) or (shutdown-agents) for future/agents, right? So same behavior regarding that. And the -X change currently is a good intermediate solution to this.

Alex Miller (Clojure team) 2025-09-02T18:48:15.482949Z

Yes, or alternately since the -X stub goes through main, could handle both there

Alex Miller (Clojure team) 2025-09-02T18:48:48.703229Z

Tbd whether changing main or clojure behavior is viable w/o being a breaking change

wouldn't it be nice to have this the same without going through main, e.g. uberjars?

anyway, thanks for the info, I think adopting this behavior in bb is a good decision (now it's completely on the other side, ignoring non-daemon threads)

Alex Miller (Clojure team) 2025-09-02T18:49:48.482209Z

Yes, it would be nice :)

🙌 1

dang. all tests passing, but only the linux static binaries are stuck on:

lein test babashka.agent-test
=== agent-binding-conveyance-test
which is this test program:
(def ^:dynamic *foo* 1) (def a (agent nil)) (binding [*foo* 2] (send-off a (fn [_] *foo*))) (await a) @a
😓

Hmm, I narrowed this down even more now. With a clojure -X invocation I can get a non-exiting program using agents.

(ns repro)

(def ^:dynamic *foo* 1)
(def a (agent nil))

(defn -main [_]
  (prn (-> (java.lang.ProcessHandle/current) (.pid)))
  (binding [*foo* 2]
    (send-off a (fn [_] *foo*)))
  (await a)
  (prn @a))
$ clj -Sdeps '{:paths ["."]}' -X repro/-main
81927
2
and then it hangs. Is this intentional @alexmiller?

Alex Miller (Clojure team) 2025-09-04T13:06:51.313839Z

Hangs forever or 1 min? If you thread dump, what’s it doing?

maybe 1 minute, let me wait longer and also thread dump

waited for a couple of minutes, still hanging. thread dump: https://gist.github.com/borkdude/6535c668c780eeed8a315e4d512909f1

> "clojure-agent-send-pool-0" #35 [40707] prio=5 os_prio=31 cpu=0,08ms elapsed=302,20s tid=0x0000000101a0fa00 nid=40707 waiting on condition

btw the binding isn't relevant, this is a smaller repro:

(ns repro)

(defn -main [_]
  (let [a (agent nil)]
    (prn (-> (java.lang.ProcessHandle/current) (.pid)))
    (send-off a (fn [_] 3))
    (await a)
    (prn @a)))

interestingly if I remove the await it does finish

smaller repro:

(ns repro)

(defn -main [_]
  (let [a (agent nil)]
    (prn (-> (java.lang.ProcessHandle/current) (.pid)))
    (send a (fn [_] 3))))

seems like send triggers it

also made a GraalVM native-image report since their behavior isn't according to the JVM except on linux + musl https://github.com/oracle/graal/issues/12116

Let me know if I should make an ask about this