All posts
7 min readDavid Nussio

Testing a keychain CLI on macOS, Linux and Windows in CI

If your tool stores secrets in the operating system's credential store, the code you most need to test is the code that talks to a system service. CI runners don't give you that service ready to use, and on your own laptop you don't want a test suite writing into your real keychain.

envsec keeps secret values in the macOS Keychain, in the Linux Secret Service (through secret-tool) and in the Windows Credential Manager, and the metadata about them in a SQLite file. Each of those three adapters shells out to a different program with different quoting rules and different failure modes. This post describes how the test suite reaches all three on GitHub Actions, what I had to set up on each runner, and the failures that taught me the most.

Three layers of tests

The tests are split by how much of the real system they touch:

The last two are one script with two modes. Without ENVSEC_E2E_CLI, it runs an alternative entry point that replaces only the credential store with a JSON file. With it, it runs the real CLI:

bash
if [[ -n "${ENVSEC_E2E_CLI:-}" ]]; then
  CLI="$ENVSEC_E2E_CLI"
  export ENVSEC_E2E_ISOLATED="${ENVSEC_E2E_ISOLATED:-0}"
else
  CLI="$SCRIPT_DIR/e2e-main.mjs"
  export ENVSEC_E2E_ISOLATED=1
fi

TMPDIR_TEST=$(mktemp -d)
export ENVSEC_DB="$TMPDIR_TEST/store.sqlite"
export ENVSEC_E2E_KEYCHAIN="$TMPDIR_TEST/keychain.json"
trap 'cleanup_secrets; rm -rf "$TMPDIR_TEST"' EXIT

cleanup_secrets

Because envsec is built from Effect services, the isolated entry point doesn't mock anything clever. It builds the same SecretStore layer the CLI uses, with the real SQLite metadata store and a file-backed KeychainAccess, and hands it to the same runner:

ts
// packages/cli/test/e2e-main.mjs (trimmed)
const fileKeychainLayer = Layer.succeed(
  KeychainAccess,
  KeychainAccess.of({ get, remove, set }) // backed by a JSON file
);

const metadataLayer = SqliteMetadataStoreLive.pipe(
  Layer.provide(DatabaseConfigFrom(databasePath))
);
const secretStoreLayer = SecretStore.layerNoDeps.pipe(
  Layer.provide(Layer.merge(fileKeychainLayer, metadataLayer))
);

runCliWithLayer(cachePath, secretStoreLayer);

So the default pnpm --filter envsec test never touches the keychain on my machine. Exercising the native adapter is an explicit opt-in:

Terminal
# Isolated: JSON-file keychain, temporary database
$pnpm --filter envsec test
# Native: the real credential store of this machine
$ENVSEC_E2E_CLI="$PWD/packages/cli/dist/main.js" \
$ ENVSEC_E2E_ISOLATED=0 \
$ pnpm --filter envsec test

Isolating the metadata database

The SQLite file is the easy half. Every run points ENVSEC_DB at a fresh mktemp -d directory and a trap removes it on exit, so a test can never read or damage the real ~/.envsec/store.sqlite.

The credential store is shared, though. A temporary database doesn't give you a temporary keychain, which is why every test secret lives under a dedicated test.e2e* context, and why cleanup_secrets runs both before the first test and from the exit trap. A run that crashed halfway leaves entries behind; the next run deletes them before it starts.

Isolation is also only as good as every line that touches the variable. The Windows script once tested a custom database path and then reset $env:ENVSEC_DB to $null. Everything after that point quietly fell back to the default database in the runner's home directory. It surfaced as a Windows-only failure, and the fix was to restore the primary temporary path instead of clearing it.

Giving each runner a credential store

macOS

Nothing to set up. The workflow has no keychain step for macOS: the tests call the CLI, which calls security add-generic-password and friends against the runner user's default keychain.

Linux

secret-tool talks to a Secret Service provider over the D-Bus session bus. A headless Ubuntu runner doesn't come with either one running, so the workflow installs GNOME Keyring and starts both:

yaml
- name: Install gnome-keyring (Linux)
  if: runner.os == 'Linux'
  run: |
    sudo apt-get update
    sudo apt-get install -y gnome-keyring dbus-x11 \
      libsecret-1-0 libsecret-tools

- name: Start D-Bus & unlock keyring (Linux)
  if: runner.os == 'Linux'
  run: |
    eval "$(dbus-launch --sh-syntax)"
    echo "DBUS_SESSION_BUS_ADDRESS=$DBUS_SESSION_BUS_ADDRESS" >> "$GITHUB_ENV"
    echo "test" | gnome-keyring-daemon --unlock --components=secrets

- name: Run E2E tests
  run: bash packages/cli/test/e2e-test.sh
  env:
    ENVSEC_E2E_CLI: ${{ github.workspace }}/packages/cli/dist/main.js
    ENVSEC_E2E_ISOLATED: "0"

Three details matter here:

Windows

Also nothing to set up: the Credential Manager is there for the runner user. envsec reaches it from PowerShell by calling the Win32 CredWriteW, CredReadW and CredDeleteW functions through P/Invoke.

One test, three native commands

The test I find most useful deletes a secret behind envsec's back, leaving its metadata orphaned, and then checks that get fails with a message that suggests envsec delete, and that delete cleans up. Simulating an out-of-band deletion needs each platform's own tool:

bash
# Delete directly from the credential store (bypass envsec),
# leaving metadata orphaned.
if [[ "$ENVSEC_E2E_ISOLATED" == "1" ]]; then
  node "$CLI" __e2e_delete_keychain "envsec.${CTX_STALE}.stale" "secret"
elif [[ "$(uname)" == "Darwin" ]]; then
  security delete-generic-password \
    -s "envsec.${CTX_STALE}.stale" -a "secret"
else
  secret-tool clear \
    service "envsec.${CTX_STALE}.stale" account "secret"
fi

And on Windows, from PowerShell:

powershell
& cmd /c "cmdkey /delete:`"envsec:envsec.${CTX_STALE}.stale/secret`""

What broke, and why

Most of what I learned came from differences between operating systems rather than from regressions. A few examples:

Shell quoting on Windows

The first Windows adapter shelled out to cmdkey through nested shells. A value like p@ss w0rd!#$% did not survive the trip, and my first reaction was to change the Windows test value to p@ssw0rd_S3cr3t so it would pass. That made the test green and the bug invisible. The real fix came a day later: the adapter moved to P/Invoke, and the test went back to the same special characters the Unix script uses.

Non-ASCII values on macOS

An emoji or an accented character came back from security as hex. Rather than decode per platform, envsec now base64-encodes every value before handing it to the credential store, with an envsec:b64: prefix, and decodes it on the way out. The E2E suite stores hello ⭐ world 🚀 and café résumé naïve on all three systems to keep it that way.

The test tools themselves

One macOS failure had nothing to do with envsec: a check used grep -P, which the BSD grep on macOS doesn't support. On Windows, the GPG share tests are skipped: the gpg on the runner is an MSYS2 build that rewrites GNUPGHOME from a Windows path into a broken Unix-style one. Sharing is tested on macOS and Linux only.

The doctor hang on Linux

When I added envsec doctor, the Ubuntu E2E job stopped finishing. One run sat for more than 13 minutes before it was cancelled. I skipped the doctor tests on Linux CI and moved on:

bash
if [[ "$ENVSEC_E2E_ISOLATED" == "1" ]]; then
  echo "  ⚠ Skipping doctor tests with the isolated credential-store fixture"
elif [[ "$OSTYPE" == "linux-gnu"* ]] && [[ -n "${CI:-}" ]]; then
  echo "  ⚠ Skipping doctor tests on Linux CI (known to hang in GitHub Actions)"
else
  # ... doctor assertions
fi

That was a skip, not a fix. Rereading the code for this post, I found the cause. The Linux branch of the read/write probe ran secret-tool store, which reads the secret from stdin when stdin isn't a terminal, but never wrote to stdin or closed it:

ts
// packages/cli/src/cli/doctor.ts, Linux branch (trimmed)
const setR = await exec("secret-tool", [
  "store", "--label", "envsec doctor probe",
  "service", testService, "account", testAccount,
]);
// secret-tool store reads from stdin — we can't easily pipe here,
// so just check if the tool is callable

Node's execFile gives the child an open stdin pipe, so a program that reads until end of file waits forever. The real Linux adapter does write the value and close stdin, which is why the rest of the Linux suite never hung. So the hang wasn't specific to CI, and the honest fix was in doctor, not in the test script. In envsec 1.1.3 the probe writes its value, closes stdin, reads the value back and clears the item, like the macOS branch always did:

ts
// 1.1.3: stdin is always closed, and a probe can't run forever
const running = execFileAsync(cmd, args, { timeout: EXEC_TIMEOUT_MS });
running.child.stdin?.on("error", ignoreStreamError);
running.child.stdin?.end(stdin);

// Linux branch: store, read back, clear
const setR = await exec(
  "secret-tool",
  ["store", "--label", "envsec doctor probe", ...attributes],
  testValue
);
const getR = await exec("secret-tool", ["lookup", ...attributes]);
await exec("secret-tool", ["clear", ...attributes]);

The skip is gone, and the Linux job now checks that doctor reports a full write, read and delete. The error handler matters too: if the child exits before reading stdin, writing to the pipe raises EPIPE, and without a listener that crashes Node instead of failing the check. While I was there, the Windows check stopped looking for cmdkey, which the adapter no longer uses, and checks for the PowerShell Add-Type that its P/Invoke code needs.

The broader lesson is about timeouts. Neither workflow sets timeout-minutes, so a hung job runs until someone cancels it or GitHub's default job limit ends it. The CLI smoke tests do better: the test for a prompt that gets no input spawns the CLI with a 10-second timeout and asserts that it exited with an error instead of being killed. The E2E script also ends by checking that no dist/main.js process is still running.

One setup action for every job

Every job needs the same preparation: Node, pnpm, a frozen-lockfile install and a build. I moved it into a composite action so the E2E matrix, the CI checks and the release jobs can't drift apart:

yaml
# .github/actions/setup-and-build/action.yml (trimmed)
inputs:
  node-version:
    default: "24"
  registry-url:
    default: ""

runs:
  using: composite
  steps:
    - uses: actions/setup-node@v7
      with:
        node-version: ${{ inputs.node-version }}
        registry-url: ${{ inputs.registry-url || '' }}
    - uses: pnpm/action-setup@v6
    - name: Install dependencies
      shell: bash
      run: pnpm install --frozen-lockfile
    - name: Build
      shell: bash
      run: pnpm run build

The registry-url input exists for the npm publish job; the default Node version moved from 22 to 24 in one place. The trade-off is that pnpm run build runs turbo run build for the whole monorepo, website included, even in jobs that only need the CLI.

The workflows that use it are split by cost:

The Ubuntu runners are pinned to ubuntu-24.04 rather than ubuntu-latest, because the keyring packages above are exactly the kind of thing a new image can break without notice.

What this does not cover

In short

Run the same end-to-end script in two modes: a file-backed keychain for day-to-day work, the real one in CI. Give every run its own database, keep test secrets in their own contexts, and clean them before and after. On Linux, start a D-Bus session and unlock GNOME Keyring yourself; macOS and Windows need nothing. And when a platform-specific test fails, be suspicious of the fix that only changes the test. The adapters themselves are described in One CLI, three keychains, and the workflows are in the envsec repository.