Testing a keychain CLI on macOS, Linux and Windows in CI
If your tool stores secrets in the operating system's credential store, the code you most need to test is the code that talks to a system service. CI runners don't give you that service ready to use, and on your own laptop you don't want a test suite writing into your real keychain.
envsec keeps secret values in the macOS Keychain, in the Linux Secret Service (through secret-tool) and in the Windows Credential Manager, and the metadata about them in a SQLite file. Each of those three adapters shells out to a different program with different quoting rules and different failure modes. This post describes how the test suite reaches all three on GitHub Actions, what I had to set up on each runner, and the failures that taught me the most.
Three layers of tests
The tests are split by how much of the real system they touch:
- Unit and contract tests run with
node --testincore,sdkandcli. The core tests provide an in-memoryKeychainAccesslayer (aMap) and a real SQLite file in a temporary directory. The CLI smoke tests spawn the builtdist/main.jswith--dbpointing at a temporary file and stick to commands that never reach the keychain. - The end-to-end script,
packages/cli/test/e2e-test.sh, walks through the whole CLI: add, get, list, search, env files, load, delete, run, saved commands, expiry and audit, GPG sharing, rename, move, copy, doctor, completions, shell, secret generation and rescue.e2e-test.ps1mirrors it for Windows. - The same script against the real credential store, which is what CI runs on macOS, Ubuntu and Windows.
The last two are one script with two modes. Without ENVSEC_E2E_CLI, it runs an alternative entry point that replaces only the credential store with a JSON file. With it, it runs the real CLI:
if [[ -n "${ENVSEC_E2E_CLI:-}" ]]; then
CLI="$ENVSEC_E2E_CLI"
export ENVSEC_E2E_ISOLATED="${ENVSEC_E2E_ISOLATED:-0}"
else
CLI="$SCRIPT_DIR/e2e-main.mjs"
export ENVSEC_E2E_ISOLATED=1
fi
TMPDIR_TEST=$(mktemp -d)
export ENVSEC_DB="$TMPDIR_TEST/store.sqlite"
export ENVSEC_E2E_KEYCHAIN="$TMPDIR_TEST/keychain.json"
trap 'cleanup_secrets; rm -rf "$TMPDIR_TEST"' EXIT
cleanup_secretsBecause envsec is built from Effect services, the isolated entry point doesn't mock anything clever. It builds the same SecretStore layer the CLI uses, with the real SQLite metadata store and a file-backed KeychainAccess, and hands it to the same runner:
// packages/cli/test/e2e-main.mjs (trimmed)
const fileKeychainLayer = Layer.succeed(
KeychainAccess,
KeychainAccess.of({ get, remove, set }) // backed by a JSON file
);
const metadataLayer = SqliteMetadataStoreLive.pipe(
Layer.provide(DatabaseConfigFrom(databasePath))
);
const secretStoreLayer = SecretStore.layerNoDeps.pipe(
Layer.provide(Layer.merge(fileKeychainLayer, metadataLayer))
);
runCliWithLayer(cachePath, secretStoreLayer);So the default pnpm --filter envsec test never touches the keychain on my machine. Exercising the native adapter is an explicit opt-in:
$# Isolated: JSON-file keychain, temporary database$pnpm --filter envsec test$# Native: the real credential store of this machine$ENVSEC_E2E_CLI="$PWD/packages/cli/dist/main.js" \$ ENVSEC_E2E_ISOLATED=0 \$ pnpm --filter envsec testIsolating the metadata database
The SQLite file is the easy half. Every run points ENVSEC_DB at a fresh mktemp -d directory and a trap removes it on exit, so a test can never read or damage the real ~/.envsec/store.sqlite.
The credential store is shared, though. A temporary database doesn't give you a temporary keychain, which is why every test secret lives under a dedicated test.e2e* context, and why cleanup_secrets runs both before the first test and from the exit trap. A run that crashed halfway leaves entries behind; the next run deletes them before it starts.
Isolation is also only as good as every line that touches the variable. The Windows script once tested a custom database path and then reset $env:ENVSEC_DB to $null. Everything after that point quietly fell back to the default database in the runner's home directory. It surfaced as a Windows-only failure, and the fix was to restore the primary temporary path instead of clearing it.
Giving each runner a credential store
macOS
Nothing to set up. The workflow has no keychain step for macOS: the tests call the CLI, which calls security add-generic-password and friends against the runner user's default keychain.
Linux
secret-tool talks to a Secret Service provider over the D-Bus session bus. A headless Ubuntu runner doesn't come with either one running, so the workflow installs GNOME Keyring and starts both:
- name: Install gnome-keyring (Linux)
if: runner.os == 'Linux'
run: |
sudo apt-get update
sudo apt-get install -y gnome-keyring dbus-x11 \
libsecret-1-0 libsecret-tools
- name: Start D-Bus & unlock keyring (Linux)
if: runner.os == 'Linux'
run: |
eval "$(dbus-launch --sh-syntax)"
echo "DBUS_SESSION_BUS_ADDRESS=$DBUS_SESSION_BUS_ADDRESS" >> "$GITHUB_ENV"
echo "test" | gnome-keyring-daemon --unlock --components=secrets
- name: Run E2E tests
run: bash packages/cli/test/e2e-test.sh
env:
ENVSEC_E2E_CLI: ${{ github.workspace }}/packages/cli/dist/main.js
ENVSEC_E2E_ISOLATED: "0"Three details matter here:
dbus-launch --sh-syntaxprints shell assignments for the new bus. Each GitHub Actions step is a new shell, so the address is written to$GITHUB_ENV, where the later test step picks it up.gnome-keyring-daemon --unlockreads a password from stdin and uses it to unlock the login keyring, creating it if it doesn't exist. The password istestbecause this keyring lives for one job.--components=secretsstarts only the Secret Service part, not the SSH or PKCS#11 agents.
Windows
Also nothing to set up: the Credential Manager is there for the runner user. envsec reaches it from PowerShell by calling the Win32 CredWriteW, CredReadW and CredDeleteW functions through P/Invoke.
One test, three native commands
The test I find most useful deletes a secret behind envsec's back, leaving its metadata orphaned, and then checks that get fails with a message that suggests envsec delete, and that delete cleans up. Simulating an out-of-band deletion needs each platform's own tool:
# Delete directly from the credential store (bypass envsec),
# leaving metadata orphaned.
if [[ "$ENVSEC_E2E_ISOLATED" == "1" ]]; then
node "$CLI" __e2e_delete_keychain "envsec.${CTX_STALE}.stale" "secret"
elif [[ "$(uname)" == "Darwin" ]]; then
security delete-generic-password \
-s "envsec.${CTX_STALE}.stale" -a "secret"
else
secret-tool clear \
service "envsec.${CTX_STALE}.stale" account "secret"
fiAnd on Windows, from PowerShell:
& cmd /c "cmdkey /delete:`"envsec:envsec.${CTX_STALE}.stale/secret`""What broke, and why
Most of what I learned came from differences between operating systems rather than from regressions. A few examples:
Shell quoting on Windows
The first Windows adapter shelled out to cmdkey through nested shells. A value like p@ss w0rd!#$% did not survive the trip, and my first reaction was to change the Windows test value to p@ssw0rd_S3cr3t so it would pass. That made the test green and the bug invisible. The real fix came a day later: the adapter moved to P/Invoke, and the test went back to the same special characters the Unix script uses.
Non-ASCII values on macOS
An emoji or an accented character came back from security as hex. Rather than decode per platform, envsec now base64-encodes every value before handing it to the credential store, with an envsec:b64: prefix, and decodes it on the way out. The E2E suite stores hello ⭐ world 🚀 and café résumé naïve on all three systems to keep it that way.
The test tools themselves
One macOS failure had nothing to do with envsec: a check used grep -P, which the BSD grep on macOS doesn't support. On Windows, the GPG share tests are skipped: the gpg on the runner is an MSYS2 build that rewrites GNUPGHOME from a Windows path into a broken Unix-style one. Sharing is tested on macOS and Linux only.
The doctor hang on Linux
When I added envsec doctor, the Ubuntu E2E job stopped finishing. One run sat for more than 13 minutes before it was cancelled. I skipped the doctor tests on Linux CI and moved on:
if [[ "$ENVSEC_E2E_ISOLATED" == "1" ]]; then
echo " ⚠ Skipping doctor tests with the isolated credential-store fixture"
elif [[ "$OSTYPE" == "linux-gnu"* ]] && [[ -n "${CI:-}" ]]; then
echo " ⚠ Skipping doctor tests on Linux CI (known to hang in GitHub Actions)"
else
# ... doctor assertions
fiThat was a skip, not a fix. Rereading the code for this post, I found the cause. The Linux branch of the read/write probe ran secret-tool store, which reads the secret from stdin when stdin isn't a terminal, but never wrote to stdin or closed it:
// packages/cli/src/cli/doctor.ts, Linux branch (trimmed)
const setR = await exec("secret-tool", [
"store", "--label", "envsec doctor probe",
"service", testService, "account", testAccount,
]);
// secret-tool store reads from stdin — we can't easily pipe here,
// so just check if the tool is callableNode's execFile gives the child an open stdin pipe, so a program that reads until end of file waits forever. The real Linux adapter does write the value and close stdin, which is why the rest of the Linux suite never hung. So the hang wasn't specific to CI, and the honest fix was in doctor, not in the test script. In envsec 1.1.3 the probe writes its value, closes stdin, reads the value back and clears the item, like the macOS branch always did:
// 1.1.3: stdin is always closed, and a probe can't run forever
const running = execFileAsync(cmd, args, { timeout: EXEC_TIMEOUT_MS });
running.child.stdin?.on("error", ignoreStreamError);
running.child.stdin?.end(stdin);
// Linux branch: store, read back, clear
const setR = await exec(
"secret-tool",
["store", "--label", "envsec doctor probe", ...attributes],
testValue
);
const getR = await exec("secret-tool", ["lookup", ...attributes]);
await exec("secret-tool", ["clear", ...attributes]);The skip is gone, and the Linux job now checks that doctor reports a full write, read and delete. The error handler matters too: if the child exits before reading stdin, writing to the pipe raises EPIPE, and without a listener that crashes Node instead of failing the check. While I was there, the Windows check stopped looking for cmdkey, which the adapter no longer uses, and checks for the PowerShell Add-Type that its P/Invoke code needs.
The broader lesson is about timeouts. Neither workflow sets timeout-minutes, so a hung job runs until someone cancels it or GitHub's default job limit ends it. The CLI smoke tests do better: the test for a prompt that gets no input spawns the CLI with a 10-second timeout and asserts that it exited with an error instead of being killed. The E2E script also ends by checking that no dist/main.js process is still running.
One setup action for every job
Every job needs the same preparation: Node, pnpm, a frozen-lockfile install and a build. I moved it into a composite action so the E2E matrix, the CI checks and the release jobs can't drift apart:
# .github/actions/setup-and-build/action.yml (trimmed)
inputs:
node-version:
default: "24"
registry-url:
default: ""
runs:
using: composite
steps:
- uses: actions/setup-node@v7
with:
node-version: ${{ inputs.node-version }}
registry-url: ${{ inputs.registry-url || '' }}
- uses: pnpm/action-setup@v6
- name: Install dependencies
shell: bash
run: pnpm install --frozen-lockfile
- name: Build
shell: bash
run: pnpm run buildThe registry-url input exists for the npm publish job; the default Node version moved from 22 to 24 in one place. The trade-off is that pnpm run build runs turbo run build for the whole monorepo, website included, even in jobs that only need the CLI.
The workflows that use it are split by cost:
ci.ymlruns on every push and pull request tomainandbeta, on Ubuntu only: Ultracite (Oxlint and Oxfmt), typecheck, unit tests, and a compiled-binary smoke test withnode_modulesdeleted.e2e.ymlruns the native suite onmacos-latest,ubuntu-24.04andwindows-latest, only when something underpackages/or the workspace config changes. The matrix setsfail-fast: false, so a Linux failure doesn't cancel the macOS run that would have told me whether the bug is platform-specific.
The Ubuntu runners are pinned to ubuntu-24.04 rather than ubuntu-latest, because the keyring packages above are exactly the kind of thing a new image can break without notice.
What this does not cover
- The E2E suite runs
dist/main.json Node. The standalone Bun binaries get smoke tests that avoid the keychain, not the full suite (more on that in Shipping one CLI to Homebrew, npm, mise and a standalone binary). - Linux is tested against GNOME Keyring only, not KWallet or other Secret Service providers.
- The keychains in CI are always unlocked and never show a permission dialog, so prompts and locked-keychain errors are not covered by CI.
shareis untested on Windows, anddoctorwas untested on Linux until 1.1.3.
In short
Run the same end-to-end script in two modes: a file-backed keychain for day-to-day work, the real one in CI. Give every run its own database, keep test secrets in their own contexts, and clean them before and after. On Linux, start a D-Bus session and unlock GNOME Keyring yourself; macOS and Windows need nothing. And when a platform-specific test fails, be suspicious of the fix that only changes the test. The adapters themselves are described in One CLI, three keychains, and the workflows are in the envsec repository.