agent-testing
Agent Testing is a local-first end-to-end verification framework for testing agentic systems across multiple surfaces (CLI, web, Electron) following a four-step contract: environment setup with dependency validation, authentication verification, test execution, and structured reporting. Use this skill when building comprehensive test suites that require consistent validation across standalone applications and multiple deployment targets while ensuring dependencies and authentication are healthy before test execution begins.
git clone --depth 1 https://github.com/lobehub/lobehub /tmp/agent-testing && cp -r /tmp/agent-testing/.agents/skills/agent-testing ~/.claude/skills/agent-testingSKILL.md
# Agent Testing (Agentic End-to-End Verification) One skill for all agentic end-to-end testing — local-first today, designed to also run as full cloud automation. Every test session follows the same contract: ``` Step -1: Plan approval → Step 0: Env + Auth → Step 1: Pick surface → Step 2: Run → Step 3: Structured report → Step 4: Publish to LobeHub ``` ## Step -1 — Plan approval for non-trivial tests Skip directly to Step 0 if: the test is a single re-run after a fix, the plan was already agreed on, or the user gave exact commands. Otherwise, propose a test plan (surface, cases, expected evidence, assumptions) and use the runtime structured question tool (`request_user_input` / ask-user-question equivalent) with two fixed choices: 1. `开始执行 (Recommended)` — 测试方案没问题,开始执行 2. `先讨论下` — 方案有问题,先讨论下 Wait for the user's choice before proceeding. ## Step 0 — Environment setup + auth check (mandatory) Step 0 is about getting the environment ready: **dependencies are healthy** and **auth is green**. A test run that dies halfway on a missing dependency or a login wall wastes the whole session — clear both gates BEFORE writing a single test step. ### 0.0 Resolve the current test environment Before starting a dev server, checking auth, opening agent-browser, or writing test steps, print and confirm the current local test environment: ```bash ./.agents/skills/agent-testing/scripts/test-env.sh ``` This command is the source of truth for local test ports. It reads the current shell plus `.env` files using the same precedence as `scripts/runWithEnv.mts`, then prints: - `APP_URL` - `PORT` - `SERVER_URL` - `AUTH_TRUSTED_ORIGINS` - `SPA_PORT` - `MOBILE_SPA_PORT` - `DESKTOP_PORT` For commands that need these values, export them from the same resolver: ```bash eval "$(./.agents/skills/agent-testing/scripts/test-env.sh --exports)" ``` Do not rely on hard-coded port tables. If the printed values do not match the running dev server, fix/export the env first, then continue. ### 0.1 Dependencies are installed — root AND standalone apps The root pnpm workspace does **NOT** cover every app: `pnpm-workspace.yaml` lists `packages/**`, `e2e`, `apps/server`, and only `apps/desktop/src/main` — **`apps/desktop` and `apps/cli` are standalone**, each keeping its own `node_modules` with its own links into `packages/`. A root install does not refresh them, so install in every app the test will touch: ```bash pnpm install # root workspace cd apps/desktop && pnpm install # Electron surface cd apps/cli && pnpm install # CLI surface ``` Symptom of a stale standalone install: the build/launch fails to resolve a recently added workspace package — `Rolldown failed to resolve import "@lobechat/<pkg>"` (Electron) or `Cannot find module '@lobechat/<pkg>'` (CLI). ### 0.2 Run scripts from the repo root All paths in this skill (`./.agents/skills/agent-testing/...`) are repo-root-relative, and background commands inherit the current working directory — a script launched while `cwd` is `apps/desktop` fails with `No such file or directory`. Verify `pwd` is the repo root before launching long-running scripts. ### 0.3 Init local dev env without `.env` For Web smoke against local code, start a **normal local dev environment**. First check the repo root for `.env`: - If `.env` exists, use the existing local configuration and start the dev server normally. - If `.env` does not exist, use the agent-testing env bootstrap. Do not start the standalone e2e server as the product under test. Use `scripts/init-dev-env.sh`. It follows the e2e setup pattern — Postgres, Redis, migrations, auth/key-vault/S3 test env, seed user — but it is owned by this skill and starts the repo's dev server (`pnpm run dev:next` / `bun run dev`), not `e2e/scripts/setup.ts --start`. The script hard-blocks when root `.env` exists, so it cannot accidentally override a user's local config. When `.env` exists, do not call any `init-dev-env.sh` subcommand. Decision flow: ```bash if [[ -f .env ]]; then bun run dev else ./.agents/skills/agent-testing/scripts/init-dev-env.sh setup-db ./.agents/skills/agent-testing/scripts/init-dev-env.sh seed-user ./.agents/skills/agent-testing/scripts/init-dev-env.sh dev fi ``` Bootstrap flow when no `.env` exists: ```bash # From repo root. Managed Postgres/Redis flow requires Docker Desktop. ./.agents/skills/agent-testing/scripts/init-dev-env.sh setup-db ./.agents/skills/agent-testing/scripts/init-dev-env.sh seed-user ./.agents/skills/agent-testing/scripts/init-dev-env.sh dev ``` If using an existing Postgres instead of the managed Docker DB, set `DATABASE_URL` and `REDIS_URL`, then skip `setup-db`: ```bash DATABASE_URL=postgresql://... REDIS_URL=redis://... ./.agents/skills/agent-testing/scripts/init-dev-env.sh migrate DATABASE_URL=postgresql://... REDIS_URL=redis://... ./.agents/skills/agent-testing/scripts/init-dev-env.sh seed-user DATABASE_URL=postgresql://... REDIS_URL=redis://... ./.agents/skills/agent-testing/scripts/init-dev-env.sh dev ``` For backend-only checks, `dev-next` is available, but Web smoke needs the full-stack `dev` command so Next can proxy the SPA HTML from Vite: ```bash ./.agents/skills/agent-testing/scripts/init-dev-env.sh dev-next ``` Useful subcommands: ```bash ./.agents/skills/agent-testing/scripts/init-dev-env.sh env # print exports ./.agents/skills/agent-testing/scripts/init-dev-env.sh write # write .records/env/agent-testing-dev.env ./.agents/skills/agent-testing/scripts/init-dev-env.sh migrate # migrations only ./.agents/skills/agent-testing/scripts/init-dev-env.sh seed-user # seed user + CLI API key ./.agents/skills/agent-testing/scripts/init-dev-env.sh qstash # local QStash for workflow paths ./.agents/skills/agent-testing/scripts/init-dev-env.sh clean-db # remove managed DB container ``` Default script env: - `APP_URL=http://localhost:3010` - `DATABASE_URL=postgresql://postgres:postgres@localhost:5433/postgres` - `DATA
Add documentation for a new AI provider — usage docs, env vars, Docker config, image resources.
Add server-side environment variables that control default values for user settings.
Agent runtime lifecycle hooks. Use for before/after tool or step hooks, tool mocks, human intervention, sub-agent calls, context compression, evals, tracing, callAgent, or lifecycle events.
Build or extend LobeHub Agent Signal pipelines. Use for signal sources, signal/action types, policies, middleware, workflow handoff, dedupe, scope behavior, or observability.
Agent tracing CLI for execution snapshots. Use for agent-tracing, traces, snapshots, LLM call inspection, context engine data, agent step analysis, or execution debugging.
Build LobeHub builtin tool packages. Use when adding agent-callable tools, manifests, executors, runtimes, inspectors, renders, placeholders, streaming, interventions, portals, or tool registries.
Build multi-platform chat bots with the chat SDK. Use for Slack, Teams, Google Chat, Discord, GitHub, Linear bots, webhooks, mentions, slash commands, cards, modals, or streaming responses.
>