Learning Record learning from practice

· linux

Running llama.cpp as a systemd Service on Debian

Overview

systemd can start a long-running program at boot, run it under a dedicated user, restart it after an unexpected failure, and route its output to the system journal or a file. A service unit is a plain-text configuration file that describes that program and its relationships to other units.

This post adapts a llama-server example for Debian Trixie. It also explains why paths such as ~/Models/model.gguf should not be copied literally into a system service.

A service unit for llama-server

Save the following as /etc/systemd/system/llama-server.service:

[Unit]
Description=llama.cpp Qwen3.8-27B API Server
After=network.target

[Service]
Type=exec
User=llama
Group=llama
WorkingDirectory=/home/llama/llama.cpp-latest/build-intel/bin
ExecStart=/home/llama/llama.cpp-latest/build-intel/bin/llama-server \
  --model /home/llama/Models/Qwen/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf \
  --host 0.0.0.0 \
  --port 11435 \
  --ctx-size 32768 \
  --threads 128 \
  --threads-batch 128 \
  --parallel 4 \
  --jinja
Restart=on-failure
RestartSec=30
Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
Environment="LD_LIBRARY_PATH=/home/llama/llama.cpp-latest/build-intel/bin:/opt/intel/oneapi/mkl/2025.1/lib:/opt/intel/oneapi/compiler/2025.1/lib"
LogsDirectory=llama-server
LogsDirectoryMode=0750
StandardOutput=append:/var/log/llama-server/llama.log
StandardError=inherit

[Install]
WantedBy=multi-user.target

The paths above assume that the llama account’s home directory is /home/llama and that the build, model, and Intel oneAPI libraries are at those locations. Change them to match the actual machine. The llama user must be able to execute the server binary and read the model and shared libraries.

Why use absolute paths?

In a shell, ~ is shorthand that the shell expands to a home directory. systemd does not run ExecStart= as a shell command, so shell syntax such as tilde expansion is not available there. The Environment= values are also unit-file values, not shell commands. Use complete paths in ExecStart=, WorkingDirectory=, and environment values instead.

The trailing backslashes in ExecStart= are valid systemd line continuations; they join the following line to the current one. systemd parses the command and its arguments itself. It does not process shell operators, pipes, or redirections. If you need shell behavior, invoke a shell explicitly and quote its command carefully; this service does not need one.

What the sections and directives mean

[Unit]: description and ordering

Description= gives the service a human-readable label. After=network.target sets start/stop ordering relative to network.target; it does not by itself pull in another unit or guarantee that a remote server is reachable.

[Service]: process configuration

  • Type=exec tells systemd to consider startup successful once it has successfully executed the server binary. This catches setup failures such as a missing executable or service user, but does not mean the model has finished loading or that the API is ready to accept requests.
  • User= and Group= run the service with the dedicated llama account rather than as root.
  • WorkingDirectory= sets the process’s working directory. It must be an absolute path that exists when the service starts.
  • ExecStart= runs the server with the supplied model, host, port, context, thread, parallelism, and Jinja options.
  • Restart=on-failure asks systemd to restart the service after an unsuccessful exit or signal. RestartSec=30 adds a 30-second delay before an automatic restart. Restart=always is another option, but also restarts after a clean exit, which is not usually desirable for a server process.
  • Environment= sets variables for the process. Keep LD_LIBRARY_PATH only if the server needs those non-standard library directories at runtime.
  • LogsDirectory=llama-server asks systemd to create /var/log/llama-server for this service and set its ownership to the configured service user and group. LogsDirectoryMode=0750 limits directory access to that user, its group, and root.
  • Keep this directory dedicated to the service. If an existing log directory has different ownership, systemd may adjust ownership of its contents to match User= and Group=.
  • StandardOutput=append:... appends standard output to the named file. StandardError=inherit sends standard error to the same destination.

The --host 0.0.0.0 option is an important deployment choice: it binds the server to all local IPv4 interfaces, not just loopback. Only use it when remote clients should be able to connect, and restrict access with suitable firewall rules and application or reverse-proxy authentication.

The append destination does not define a log-retention policy. Add a separate rotation policy if the log must be kept within a size or age limit.

Choose where the server logs go

The example above writes the server’s standard output to a file, which is useful when you want a dedicated log at a predictable path. You can instead let systemd send the server output to the journal. In the [Service] section, replace the file-logging directives:

LogsDirectory=llama-server
LogsDirectoryMode=0750
StandardOutput=append:/var/log/llama-server/llama.log
StandardError=inherit

with:

StandardOutput=journal
StandardError=inherit

With the journal choice, remove LogsDirectory= and LogsDirectoryMode= because the service no longer writes to that log directory. StandardError=inherit sends standard error to the same destination as standard output. Read or follow this service’s journal entries with:

sudo journalctl -u llama-server.service
sudo journalctl -u llama-server.service -f

Choose file output if you need a standalone log file and are prepared to manage its rotation. Choose the journal if you prefer to query the service’s logs by unit with journalctl and avoid configuring a per-service file path. The unit can use one destination at a time. After changing the directives, reload the unit definition and restart the service to apply the new output destination:

sudo systemctl daemon-reload
sudo systemctl restart llama-server.service

[Install]: boot-time enablement

WantedBy=multi-user.target defines the target that systemctl enable uses to arrange for this service to start during normal multi-user boot. The [Install] section is acted on when enabling or disabling a unit; it does not start the service by itself.

Install, start, and inspect the service

First confirm the account, executable, model, and library paths in the unit. Then create the unit file and ask systemd to load it:

sudoedit /etc/systemd/system/llama-server.service
sudo systemctl daemon-reload

Enable the service at boot and start it now:

sudo systemctl enable --now llama-server.service

Check its state and follow the application log:

sudo systemctl status llama-server.service
sudo journalctl -u llama-server.service -b --no-pager
sudo tail -f /var/log/llama-server/llama.log

The journal is useful for systemd setup errors; the application log contains the server’s standard output and error. Common causes include a typo in an absolute path, an unreadable model file, missing shared libraries, or an address/port that cannot be bound.

Debian note

The Debian Trixie systemd package provides the system and service manager, but installing that package alone does not necessarily switch the machine’s init system. Debian’s package information notes that booting with systemd as PID 1 requires booting with init=/lib/systemd/systemd or installing systemd-sysv as well.

References