NodeTool runs Python nodes and local-compute providers in a separate Python worker process. The desktop app and server spawn that worker and communicate with it over a local stdio RPC protocol.
Why this protocol exists
NodeTool uses Python for:
- Python node execution
- local ML providers such as MLX and HuggingFace
- media processing and model-specific dependencies
The TypeScript runtime uses stdio instead of a localhost socket for the worker because it is simpler to supervise, avoids port-management issues, and keeps the worker strictly parent-scoped.
Transport
- Parent process spawns:
python -m nodetool.worker --stdio stdin/stdout: binary protocol trafficstderr: worker logs and startup diagnostics
Framing
Each message is encoded as:
[4-byte big-endian payload length][MessagePack payload]
The payload is a MessagePack object with this envelope shape:
{
"type": "discover",
"request_id": "uuid-or-null",
"data": {}
}
Lifecycle
- TypeScript spawns the worker
- Python loads node packages and providers
- Python prints
NODETOOL_STDIO_READYonstderr - TypeScript sends
discover - Python responds with node metadata, protocol version, and any load errors
- TypeScript optionally requests
worker.status - Workflow execution uses
execute/execute.stream,cancel,provider.*, and (on v2+ workers)models.*messages
Message types
discover
Returns the worker’s executable node inventory.
Response data:
protocol_versionnodes— each entry may carryrequires_vram_gb(v4+): the approximate VRAM that node type’s weights need, in GiB. The JS side echoes it back onexecuteso the worker’s reclaim pass has a real number to target instead of a percentage threshold.load_errors
Example:
{
"type": "discover",
"request_id": "d1",
"data": {
"protocol_version": 1,
"nodes": [...],
"load_errors": [
{
"module": "nodetool.nodes.mlx.text_generation",
"phase": "module_import",
"error": "No module named 'mlx_lm'",
"error_type": "ModuleNotFoundError"
}
]
}
}
worker.status
Returns a structured worker health snapshot.
Response data:
protocol_versionnode_countprovider_countnamespacesload_errorstransportmax_frame_size
execute
Executes a Python node.
Request data:
node_typefieldssecretsblobs
Run identity, added in v4. Every field is optional, and a field the JS host cannot name is omitted rather than sent as null:
node_id— the graph node id. Becomes the constructed node’s_id. Before v4 the worker built every node with no id, soself._idwas""for all of them: a node callingset_model(self._id, …)registered into a single shared bucket, andrelease_nodes()had nothing meaningful to release.node_idis what makes the worker’s node → model map real.job_id— pairs the execution with thejob.start/job.endboundary.workflow_id,user_id— populate the worker’sWorkerContext, which previously had neither.requires_vram_gb— the hint the worker itself reported for this node type atdiscover.
These are extra dict entries, so the JS side sends them unconditionally
rather than behind the v4 capability gate: a pre-v4 worker reads the four keys
it knows and ignores the rest, and gating would only starve a worker that does
understand them whenever its worker.status has not landed yet.
Possible response types:
progressresulterror
execute.stream
Executes a Python node that streams results. Same request data as execute
(including the v4 identity fields), but the worker emits zero or more chunk
messages (each carrying partial outputs/blobs) followed by a terminal
result or error.
cancel
Requests cooperative cancellation for an in-flight execute or streaming provider request.
provider.*
Used for Python-only providers. The message families the TS bridge implements:
provider.listprovider.modelsprovider.generateprovider.streamprovider.text_to_imageprovider.image_to_imageprovider.ttsprovider.asrprovider.embedding
Streaming providers (provider.stream, provider.tts) emit zero or more chunk messages followed by a terminal result or error.
models.*
Worker model management (HuggingFace cache). These were introduced in bridge
protocol v2 and are gated by supportsModelManagement() — a v1 worker
simply does not expose them (see Versioning).
models.list_cached— list models cached on the worker’sHF_HOME(cache-only, no network)models.download— download a model onto the worker cache, streaming orderedprogressframes then a terminalresultmodels.delete— delete a cached model; returns whether it existedmodels.evict(v4, gated bysupportsJobLifecycle()) — drop loaded model weights. Optional request data narrows the scope:node_ids,job_id,target_vram_gb(stop once that many GiB are reclaimed). Response data is{evicted: string[], freed_vram_gb?: number}. This is the path for what only the JS side knows — the user switched workflows, another process wants the GPU, the worker is idle. Without it the worker only ever reclaims reactively, on its own thresholds. Calling it against a pre-v4 worker resolves to{evicted: []}instead of erroring, so a host asking to free memory never has to branch on the worker’s version.
job.*
The run boundary. Introduced in bridge protocol v4 and gated by
supportsJobLifecycle().
job.start— opens a run. Request data:job_id, plus optionalworkflow_id/user_id. Optional in the sense that the worker needs nojob.startto attribute an execution (everyexecutecarries its ownjob_id); it exists as the one place to do a single reclaim pass per run instead of one per node.job.end— closes a run. Same data plusreason(completed|failed|cancelled|abandoned). The job’s nodes are retired and their models become eligible for release. This is the callerrelease_nodes()never had: without it the worker’s model cache grows across runs and is only ever trimmed reactively under memory pressure.
job.end must fire on abnormal termination too — cancelled, client
disconnected, run abandoned — or the leak simply moves to the failure path.
On the JS side ExecutionSession sends it from the same finally that closes
the bridge, so completion, failure, cancellation and timeout all reach it. Both
calls are fire-and-forget and swallow their own failures: a boundary is
bookkeeping, and a job.end that fails against a worker already tearing down
must not turn a finished run into a failed one.
A host that owns a long-lived shared bridge (the WebSocket runner) and injects
its own executor resolver must pass that bridge to the session as
jobLifecycleBridge — that host’s bridgeFactory deliberately returns null, so
without it nothing closes the boundary for the exact deployment the shared
bridge exists for.
comfy.*
ComfyUI proxy. Introduced in bridge protocol v3, gated by supportsComfy()
— a worker offers these only when it fronts a co-located, loopback-only ComfyUI
server AND reports worker.status.comfy.enabled: true. Route comfy.* requests
only to such workers. The full field-level reference lives in
docs/comfy-proxy.md in nodetool-core; ComfyUI covers the nodes
that sit on top of these messages.
comfy.execute— submit an API-format workflow; streams its lifecycle as dedicatedcomfy.eventframes (see below), then a terminalresult/error. Cancel with the standardcancelframe.comfy.queue—{queue_running, queue_pending}comfy.interrupt— global stop of the running job (admin-only)comfy.cancel— best-effort per-prompt cancel ({prompt_id}); the safe user-facing cancelcomfy.upload/comfy.view— stage a file into / fetch a file from ComfyUI’s input dircomfy.object_info— full ComfyUI node catalogcomfy.system_stats/comfy.status— health/capacity;comfy.statusadds worker-level{enabled, url, reachable}comfy.free— unload models from VRAM without a cold restartcomfy.models.list/comfy.models.download/comfy.models.delete— manage model files on the worker’s persistent volume.comfy.models.downloadstreams genericprogressframes (NOTcomfy.event) then a terminalresult.
comfy.event
comfy.execute does not stream progress frames — ComfyUI’s events don’t
fit the {progress, total, message} shape. Instead it emits a dedicated frame:
{
"type": "comfy.event",
"request_id": "<the execute request's id>",
"data": { "event": "executing", "prompt_id": "p1", "node": "3" }
}
data.event is the discriminator, in emission order: queued → queue
(repeatable) → started / cached → executing (per node) → progress →
node_output → preview (only if previews: true) → completed or
cancelled. result is always the last frame.
Result, error, chunk, and progress
result
{
"type": "result",
"request_id": "e1",
"data": {
"outputs": { "text": "hello" },
"blobs": {}
}
}
error
{
"type": "error",
"request_id": "e1",
"data": {
"error": "Unknown node type: foo.Bar",
"traceback": "..."
}
}
chunk
Used by streaming providers and streaming provider-like operations.
progress
The worker forwards NodeProgress messages from Python execution as protocol-level progress events:
{
"type": "progress",
"request_id": "e1",
"data": {
"progress": 32,
"total": 100,
"message": "Downloading model"
}
}
Binary data
MessagePack allows binary payloads directly. NodeTool uses that for:
- input blobs in
execute.data.blobs - output blobs in
result.data.blobs - audio/image chunks for streaming provider APIs
Diagnostics and failure handling
The bridge now surfaces worker problems in-band:
discover.load_errorslists node import and metadata extraction failuresworker.status.load_errorsprovides the same information on demanderror.tracebackis returned for failed requests- recent
stderrlines are still retained on the TS side for debugging startup failures
This matters because metadata may exist for a Python node even if its module failed to import in the worker. In that case the worker is connected, but the node is unavailable. load_errors is the authoritative signal for that situation.
Limits and timeouts
- The worker and TS bridge enforce a maximum frame size via
NODETOOL_BRIDGE_MAX_FRAME_SIZE - The TS stdio bridge enforces a startup timeout (
startupTimeoutMs, default 20s) - Cancellation is cooperative
Versioning
The JS runtime and Python worker each report a BRIDGE_PROTOCOL_VERSION. Two
distinct numbers govern compatibility (see packages/protocol/src/bridge-protocol.ts):
BRIDGE_PROTOCOL_VERSION— the protocol the JS runtime currently speaks (presently4).MIN_BRIDGE_PROTOCOL_VERSION— the hard floor (presently1). The JS runtime rejects a worker only if it reports a protocol below this floor.
Compatibility rules:
-
Older worker, at or above the floor — connects normally. It does not fail startup. Additive features the worker predates are gated per-capability: a v1 worker connects fine and simply doesn’t expose the
models.*family (gated bysupportsModelManagement(), which requires v2+),comfy.*(gated bysupportsComfy(), which requires v3+ andcomfy.enabled), orjob.*/models.evict(gated bysupportsJobLifecycle(), which requires v4+). Workers that predate theprotocol_versionfield are treated as v1.The v4 identity fields on
executeare the exception that proves the rule: they are extra keys on an existing message, not a new message, so they are sent to every worker and ignored by the ones that predate them. Only new message types need a capability gate, because those are what a worker answers withUnknown message type. - Worker below
MIN_BRIDGE_PROTOCOL_VERSION— rejected atdiscoverwith an actionable error (reinstall the Python environment). This is the only startup-failing case, reserved for genuine wire breaks. - Newer worker protocol than the JS runtime — the runtime warns and assumes backward compatibility.
Notes for contributors
If you change the protocol:
- update the JS and Python protocol version constants together
- document the schema change here
- add or update end-to-end tests in
nodetool-core/tests/ - keep
discoverandworker.statusauthoritative for diagnostics
Related
- Architecture
- Developer Guide
nodetool-core/README.md