Skip to content

CLI Reference

kabactl [GLOBAL OPTIONS] <COMMAND>
CommandPurpose
serverRun the server in the foreground.
serviceRun the server as a background daemon, or install it.
clusterInvite, join, list and evict peers.
getDownload base-model weights and the voice bundle.
askRun inference locally or on a peer.
trainFine-tune a LoRA adapter on your memories.
lorasList adapters.
lora-importImport a third-party adapter by URL.
securityManage the local Safe Browsing database.
doctorHealth-check the installation.
self-updateReplace the binary with the latest release.
OptionDescription
--traceCrash and memory diagnostics: heap-profile stack traces (Linux), memory sampling, full panic backtraces and raised core-dump limits. Writes a bundle to <storage>/perf/<timestamp>/ and points perf/CURRENT at it. Same as KABA_TRACE=1, which the service daemon inherits.
-c, --config-path <FILE>Path to config.toml.
-v / -qRaise or lower verbosity.
-V, --versionPrint the version.
-h, --helpPrint help.

Runs the full server in the foreground: HTTPS API, SOCKS5 proxy and the cluster endpoint. This is what the client starts, and what the systemd units and the container image run.

Terminal window
kabactl server [--device <BACKEND>]
OptionDefaultDescription
--device <BACKEND>autoInference backend preference: auto, cuda (alias nvidia), rocm (hip, amd), wgpu (gpu, vulkan, metal) or cpu (ndarray). CPU is always the final fallback. KABA_BACKEND overrides this flag.
-b, --bind <ADDRESS>0.0.0.0See the caution below.
--port <PORT>28832See the caution below.

Only wgpu and cpu are compiled into a default build. Requesting cuda or rocm on a build without them logs a warning and falls through to the next backend.

Lifecycle controls for the same server, detached from your terminal.

SubcommandDescription
service startDaemonize the server. Fails if one is already running.
service stopSend SIGTERM, wait, then SIGKILL if it has not exited. The wait is KABA_STOP_GRACE_SECS.
service restartStop if running, then start. If the stop fails, no replacement is started.
service statusReport running (with PID), stale PID file, or not running.
service join <TICKET>Join a cluster using a ticket. Same as cluster join.
service installCopy this binary to the managed location and put it on PATH.
service uninstallStop the service and remove the unit, symlink and managed binary.

service install options:

OptionDescription
--systemSystem-wide install to /usr/local/bin. Requires root. Without it, the install is per-user to <storage>/bin with a symlink in ~/.local/bin on Linux.
--systemdAlso write, enable and start a systemd unit.
--forceOverwrite an existing managed binary.

service uninstall takes --system for a system-wide install. See services for the units themselves.

See cluster & mesh for the workflow.

SubcommandDescription
cluster invite [-n, --name <LABEL>] [--ttl <MINUTES>]Create a join ticket. --name defaults to node, --ttl to 15. Prints a kaba_invite_… string.
cluster join <TICKET>Join the cluster that issued the ticket.
cluster listList known peers.
cluster show <NAME_OR_NODE_ID>Print everything known about one peer.
cluster ping <NAME_OR_NODE_ID>Round-trip to a peer and print the latency.
cluster evict <NAME_OR_NODE_ID> [-y]Remove a peer. Without -y it only prints what it would do.
cluster repairRe-establish keys for peers that are known but keyless.
cluster tombstonesList evicted node IDs. An evicted node cannot rejoin while its tombstone exists.
cluster rejoin <NODE_ID>Clear a tombstone so that node can join again.

Downloads and prepares base-model files into <storage>/kaba-engine. Idempotent: files already present and valid are skipped. Runs on CPU, so it works on a machine with no GPU.

Terminal window
kabactl get # full prep: tokenizer, safetensors, training checkpoint, GGUF
kabactl get --model e2b # one GGUF for serving (fastest first install)
kabactl get --only-inference # GGUFs for every variant, no training prep
kabactl get --voice # speech-to-text and text-to-speech models
OptionDescription
--model <VARIANT>Fetch one serving GGUF: e2b (about 2.8 GB) or e4b (about 4.4 GB). Enough for ask and for serving adapters. train still needs the bare get.
--only-inferenceFetch the serving GGUF for every variant and skip the training conversion.
--projectionsAlso fetch the multimodal projector. Image and audio input use dedicated routes, so this is only needed for KABA_VISION_MODEL=base or older projection asks.
--voiceFetch the voice bundle. The server also provisions this automatically at start.

Generates an answer to a prompt and streams tokens to stdout.

Terminal window
kabactl ask "Summarize what I read about QUIC this week" --workspace you@example.com
kabactl ask "What is 2+2?" --peer gpu-tower

Local asks drive the engine in this process and expect the server to be stopped, because the engine and the vector store want exclusive access. Remote asks go to a cluster peer over kaba/infer/v1; that peer must have accept_remote_inference = true.

OptionDescription
--peer <NAME_OR_NODE_ID>Run on a cluster peer instead of locally.
--workspace <EMAIL>Ground the answer on that account’s encrypted memories. Prompts for the account password, or reads KABA_WORKSPACE_PASSWORD. Without it, the base model answers with no memory grounding.
--lora <NAME>Apply a named adapter. On a peer, omit it to use the peer’s configured adapter, or pass --lora "" to force the peer’s base model.
--base-model <VARIANT>e2b or e4b.
--max-tokens <N>Cap on generated tokens.
--temperature <T>Sampling temperature.
--repetition-penalty <P>Repetition penalty.
--max-context <N>Context window in tokens.
--gpu-layers <N>Pin how many model layers stay on the GPU. Omitted, the node fits to free VRAM. KABA_LLAMA_GPU_LAYERS on the serving node caps the request.
--turboUse the compressed working-memory cache.
--quantize <Q>q8 or q4.
--gguf <PATH>Use a specific GGUF file.
--perfPrint performance timing.

Trains a LoRA adapter from memories and writes it to kaba-engine/loras/<name>.{mpk,json}.

Terminal window
kabactl train --workspace you@example.com --last 30d --output-name october
kabactl train --workspace you@example.com --node gpu-tower

Local training needs a GPU build and expects the server to be stopped. With --node, your corpus is streamed to a peer over kaba/train/v1, progress is shown, and the finished adapter is pulled back; that peer must have accept_remote_training = true.

Which memories

OptionDescription
--workspace <EMAIL>The account whose encrypted memories to train on.
--since <DATE|DUR> / --until <DATE|DUR>Window bounds: YYYY-MM-DD, an ISO instant, or a duration such as 7d or 24h.
--last <DURATION>Shorthand for a trailing window. Omit all three to train on everything.
--min-records <N>Refuse to train on fewer records. Default from config: 200.

Where and what

OptionDescription
--node <NAME_OR_NODE_ID>Train on a cluster peer.
--output-name <NAME>Adapter name.
--voice <VOICE>How training answers are phrased: factual (default) or recall.
--skip-summariesSkip the summarisation pass.
--summarise-batch <N>, --summarise-batch-trustBatch the summarisation pass.
--perfPrint performance timing.

Hyperparameters (defaults come from kaba_engine_options)

OptionDescription
--epochs <N>Default 3.
--lr <LR>Peak learning rate. Default 1e-4.
--lora-rank <N>Default 16.
--batch-size <N>, --seq-len <N>, --grad-accum <N>, --warmup-ratio <RATIO>Batch shape and schedule.
--no-qloraTrain without quantization.
Terminal window
kabactl loras [--peer <NAME_OR_NODE_ID>] [--json]

Lists adapter names in kaba-engine/loras, or on a peer. An empty list just means nothing has been trained yet.

Terminal window
kabactl lora-import <URL> [--base-model <VARIANT>]

Downloads a community adapter from a Hugging Face repo URL or a direct .gguf / .safetensors link, converts it so kabactl can run it, and writes a manifest so it shows up in kabactl loras and in the policy editor. The raw download is kept alongside as provenance. Set HF_TOKEN for private repos.

Manages the local threat database used to screen browsing. Feeds are configured under security_options.feeds.

SubcommandDescription
security updateRebuild the Bloom filter and index from the configured feeds. Skipped when feeds are unchanged.
security status [--json]Indicator count, last update time and per-source results.
security check <VALUE> [--kind <KIND>]Test a value. --kind defaults to domain; the other kinds are url, ip, filename and sha256.
Terminal window
kabactl doctor [--json]

Reports PASS, WARN or FAIL for each resource kabactl depends on, and exits non-zero if anything failed.

CheckLooks at
config_tomlconfig.toml exists and parses.
tls_cert, tls_keyPresent, non-empty, and the key is not group or world readable.
sqlite_dbkaba.db exists and is writable.
lancedb_storememories/kaba is readable and writable.
engine_dir, tokenizer_json, gemma_gguf, loras_dirModel files and adapter count.
service_pid, service_logRunning, stale PID, or stopped.
managed_binary, kabactl_on_pathThe installed binary is executable and on PATH.
server_port, proxy_portPorts are free, or held by the running service.
update_tokenWhether KABA_UPDATE_TOKEN is set.
Terminal window
kabactl self-update [--check] [--to <TAG>] [--force] [--no-restart]

Fetches a release, picks the asset for this platform, verifies the downloaded binary reports the expected version, and atomically swaps it over the managed binary. If the service is running it is restarted unless --no-restart is given.

OptionDescription
--checkReport whether an update is available and exit.
--to <TAG>Install a specific release tag.
--forceReinstall even if already current.
--no-restartLeave the running service on the old binary.

Release assets exist for linux-x64, linux-arm64, macos-arm64 and windows-x64. For a private repository set KABA_UPDATE_TOKEN to a read-scope access token; it is sent as an Authorization header, never in a URL. KABA_UPDATE_BASE_URL points at a different release host.