mirror of
https://github.com/ksyasuda/dotfiles.git
synced 2026-08-17 06:18:26 -07:00
update
This commit is contained in:
@@ -4,4 +4,4 @@ description: A relentless interview to sharpen a plan or design.
|
|||||||
disable-model-invocation: true
|
disable-model-invocation: true
|
||||||
---
|
---
|
||||||
|
|
||||||
Run a `/grilling` session.
|
Call the Skill tool with "grilling".
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Grill Me"
|
||||||
|
short_description: "Sharpen a plan through interview"
|
||||||
|
policy:
|
||||||
|
allow_implicit_invocation: false
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
---
|
||||||
|
name: grilling
|
||||||
|
description: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.
|
||||||
|
---
|
||||||
|
|
||||||
|
Interview the user relentlessly until you reach a shared understanding. Map this as a **design tree**: every decision branches into the decisions that hang off it.
|
||||||
|
|
||||||
|
Work the tree in **rounds**. The **frontier** is every decision whose prerequisites are already settled: the questions you can ask _now_ without guessing at answers you haven't heard yet. Ask the whole frontier in one round: number each question and give your recommended answer. Then wait for the user's answers before the next round.
|
||||||
|
|
||||||
|
Each question should be formatted like so:
|
||||||
|
|
||||||
|
```
|
||||||
|
❓ **Q1** - **<question title>**: <question body, might be multiple paragraphs, including multiple choices>
|
||||||
|
|
||||||
|
➡️ <your recommended answer>
|
||||||
|
```
|
||||||
|
|
||||||
|
Each round the user answers reshapes the tree: settled decisions push the frontier outward and unblock questions that depended on them. Recompute the frontier and ask the next round. A question whose answer depends on another question still open in this round belongs to a _later_ round, not this one.
|
||||||
|
|
||||||
|
Finding _facts_ is your job, never the user's. When a frontier question needs a fact from the environment (filesystem, tools, etc.), dispatch a sub-agent to find it; don't ask the user for anything you could look up yourself. Don't block on it: a running exploration is an unsettled prerequisite, so only the questions downstream of it wait for the sub-agent to report; ask the rest of the frontier now. The _decisions_ are the user's: put each to them and wait.
|
||||||
|
|
||||||
|
The session is done when the frontier is empty: every branch of the design tree visited, nothing left silently assumed. Do not act on it until the user confirms you have reached a shared understanding.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Grilling"
|
||||||
|
short_description: "Stress-test thinking a round of questions at a time"
|
||||||
@@ -0,0 +1,123 @@
|
|||||||
|
# HTML Report Format
|
||||||
|
|
||||||
|
The architectural review is rendered as a single self-contained HTML file in the OS temp directory. Tailwind and Mermaid both come from CDNs. Mermaid handles graph-shaped diagrams reliably; hand-built divs and inline SVG handle the more editorial visuals (mass diagrams, cross-sections). Mix the two — don't lean on Mermaid for everything, it'll start to look generic.
|
||||||
|
|
||||||
|
## Scaffold
|
||||||
|
|
||||||
|
```html
|
||||||
|
<!doctype html>
|
||||||
|
<html lang="en">
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8" />
|
||||||
|
<title>Architecture review — {{repo name}}</title>
|
||||||
|
<script src="https://cdn.tailwindcss.com"></script>
|
||||||
|
<script type="module">
|
||||||
|
import mermaid from "https://cdn.jsdelivr.net/npm/mermaid@11/dist/mermaid.esm.min.mjs";
|
||||||
|
mermaid.initialize({ startOnLoad: true, theme: "neutral", securityLevel: "loose" });
|
||||||
|
</script>
|
||||||
|
<style>
|
||||||
|
/* small custom layer for things Tailwind doesn't cover cleanly:
|
||||||
|
dashed seam lines, hand-drawn-feeling arrow heads, etc. */
|
||||||
|
.seam { stroke-dasharray: 4 4; }
|
||||||
|
.leak { stroke: #dc2626; }
|
||||||
|
.deep { background: linear-gradient(135deg, #0f172a, #1e293b); }
|
||||||
|
</style>
|
||||||
|
</head>
|
||||||
|
<body class="bg-stone-50 text-slate-900 font-sans">
|
||||||
|
<main class="max-w-5xl mx-auto px-6 py-12 space-y-12">
|
||||||
|
<header>...</header>
|
||||||
|
<section id="candidates" class="space-y-10">...</section>
|
||||||
|
<section id="top-recommendation">...</section>
|
||||||
|
</main>
|
||||||
|
</body>
|
||||||
|
</html>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Header
|
||||||
|
|
||||||
|
Repo name, date, and a compact legend: solid box = module, dashed line = seam, red arrow = leakage, thick dark box = deep module. No introduction paragraph — straight into the candidates.
|
||||||
|
|
||||||
|
## Candidate card
|
||||||
|
|
||||||
|
The diagrams carry the weight. Prose is sparse, plain, and uses the glossary terms (from the `/codebase-design` skill) without ceremony.
|
||||||
|
|
||||||
|
Each candidate is one `<article>`:
|
||||||
|
|
||||||
|
- **Title** — short, names the deepening (e.g. "Collapse the Order intake pipeline").
|
||||||
|
- **Badge row** — recommendation strength (`Strong` = emerald, `Worth exploring` = amber, `Speculative` = slate), plus a tag for the dependency category (`in-process`, `local-substitutable`, `ports & adapters`, `mock`).
|
||||||
|
- **Files** — monospaced list, `font-mono text-sm`.
|
||||||
|
- **Before / After diagram** — the centrepiece. Two columns, side by side. See patterns below.
|
||||||
|
- **Problem** — one sentence. What hurts.
|
||||||
|
- **Solution** — one sentence. What changes.
|
||||||
|
- **Wins** — bullets, ≤6 words each. e.g. "Tests hit one interface", "Pricing logic stops leaking", "Delete 4 shallow wrappers".
|
||||||
|
- **ADR callout** (if applicable) — one line in an amber-tinted box.
|
||||||
|
|
||||||
|
No paragraphs of explanation. If the diagram needs a paragraph to be understood, redraw the diagram.
|
||||||
|
|
||||||
|
## Diagram patterns
|
||||||
|
|
||||||
|
Pick the pattern that fits the candidate. Mix them. Don't make every diagram look the same — variety is part of the point.
|
||||||
|
|
||||||
|
### Mermaid graph (the workhorse for dependencies / call flow)
|
||||||
|
|
||||||
|
Use a Mermaid `flowchart` or `graph` when the point is "X calls Y calls Z, and look at the mess." Wrap it in a Tailwind-styled card so it doesn't feel parachuted in. Style with classDef to colour leakage edges red and the deep module dark. Sequence diagrams work well for "before: 6 round-trips; after: 1."
|
||||||
|
|
||||||
|
```html
|
||||||
|
<div class="rounded-lg border border-slate-200 bg-white p-4">
|
||||||
|
<pre class="mermaid">
|
||||||
|
flowchart LR
|
||||||
|
A[OrderHandler] --> B[OrderValidator]
|
||||||
|
B --> C[OrderRepo]
|
||||||
|
C -.leak.-> D[PricingClient]
|
||||||
|
classDef leak stroke:#dc2626,stroke-width:2px;
|
||||||
|
class C,D leak
|
||||||
|
</pre>
|
||||||
|
</div>
|
||||||
|
```
|
||||||
|
|
||||||
|
### Hand-built boxes-and-arrows (when Mermaid's layout fights you)
|
||||||
|
|
||||||
|
Modules as `<div>`s with borders and labels. Arrows as inline SVG `<line>` or `<path>` elements positioned absolutely over a relative container. Reach for this when you want the "after" diagram to feel like one thick-bordered deep module with greyed-out internals — Mermaid won't render that with the right weight.
|
||||||
|
|
||||||
|
### Cross-section (good for layered shallowness)
|
||||||
|
|
||||||
|
Stack horizontal bands (`h-12 border-l-4`) to show layers a call passes through. Before: 6 thin layers each doing nothing. After: 1 thick band labelled with the consolidated responsibility.
|
||||||
|
|
||||||
|
### Mass diagram (good for "interface as wide as implementation")
|
||||||
|
|
||||||
|
Two rectangles per module — one for interface surface area, one for implementation. Before: interface rectangle is nearly as tall as the implementation rectangle (shallow). After: interface rectangle is short, implementation rectangle is tall (deep).
|
||||||
|
|
||||||
|
### Call-graph collapse
|
||||||
|
|
||||||
|
Before: a tree of function calls rendered as nested boxes. After: the same tree collapsed into one box, with the now-internal calls shown faded inside it.
|
||||||
|
|
||||||
|
## Style guidance
|
||||||
|
|
||||||
|
- Lean editorial, not corporate-dashboard. Generous whitespace. Serif optional for headings (`font-serif` works well with stone/slate).
|
||||||
|
- Colour sparingly: one accent (emerald or indigo) plus red for leakage and amber for warnings.
|
||||||
|
- Keep diagrams ~320px tall so before/after sits comfortably side by side without scrolling.
|
||||||
|
- Use `text-xs uppercase tracking-wider` for module labels inside diagrams — they should read as schematic, not as UI.
|
||||||
|
- The only scripts are the Tailwind CDN and the Mermaid ESM import. The report is otherwise static — no app code, no interactivity beyond Mermaid's own rendering.
|
||||||
|
|
||||||
|
## Top recommendation section
|
||||||
|
|
||||||
|
One larger card. Candidate name, one sentence on why, anchor link to its card. That's it.
|
||||||
|
|
||||||
|
## Tone
|
||||||
|
|
||||||
|
Plain English, concise — but the architectural nouns and verbs come straight from the `/codebase-design` skill. Concision is not an excuse to drift.
|
||||||
|
|
||||||
|
**Use exactly:** module, interface, implementation, depth, deep, shallow, seam, adapter, leverage, locality.
|
||||||
|
|
||||||
|
**Never substitute:** component, service, unit (for module) · API, signature (for interface) · boundary (for seam) · layer, wrapper (for module, when you mean module).
|
||||||
|
|
||||||
|
**Phrasings that fit the style:**
|
||||||
|
|
||||||
|
- "Order intake module is shallow — interface nearly matches the implementation."
|
||||||
|
- "Pricing leaks across the seam."
|
||||||
|
- "Deepen: one interface, one place to test."
|
||||||
|
- "Two adapters justify the seam: HTTP in prod, in-memory in tests."
|
||||||
|
|
||||||
|
**Wins bullets** name the gain in glossary terms: *"locality: bugs concentrate in one module"*, *"leverage: one interface, N call sites"*, *"interface shrinks; implementation absorbs the wrappers"*. Don't write *"easier to maintain"* or *"cleaner code"* — those terms aren't in the glossary and don't earn their place.
|
||||||
|
|
||||||
|
No hedging, no throat-clearing, no "it's worth noting that…". If a sentence could be a bullet, make it a bullet. If a bullet could be cut, cut it. If a term isn't in the `/codebase-design` glossary, reach for one that is before inventing a new one.
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
---
|
||||||
|
name: improve-codebase-architecture
|
||||||
|
description: Scan a codebase for deepening opportunities, present them as a visual HTML report, then grill through whichever one you pick.
|
||||||
|
disable-model-invocation: true
|
||||||
|
---
|
||||||
|
|
||||||
|
# Improve Codebase Architecture
|
||||||
|
|
||||||
|
Surface architectural friction and propose **deepening opportunities** — refactors that turn shallow modules into deep ones. The aim is testability and AI-navigability.
|
||||||
|
|
||||||
|
This command is _informed_ by the project's domain model and built on a shared design vocabulary:
|
||||||
|
|
||||||
|
- Call the Skill tool with "codebase-design" for the architecture vocabulary (**module**, **interface**, **depth**, **seam**, **adapter**, **leverage**, **locality**) and its principles (the deletion test, "the interface is the test surface", "one adapter = hypothetical seam, two = real"). Use these terms exactly in every suggestion — don't drift into "component," "service," "API," or "boundary."
|
||||||
|
- The domain language in `CONTEXT.md` gives names to good seams; ADRs in `docs/adr/` record decisions this command should not re-litigate.
|
||||||
|
|
||||||
|
## Process
|
||||||
|
|
||||||
|
### 1. Explore
|
||||||
|
|
||||||
|
**Scope before you scan — YAGNI.** Deepening a module pays off by making future changes to it easier, so put extra weight on the parts of the codebase that have recently changed. Decide *where* to look before you look:
|
||||||
|
|
||||||
|
- If the user named a direction — a module, a subsystem, a pain point — take it, and skip the inference below.
|
||||||
|
- Otherwise, walk back a good stretch of the commit history (`git log --oneline`) to find the codebase's hot spots — the files and areas that keep coming up — and let those paths pull your attention first. If the changes are scattered with no clear hot spot, widen the net.
|
||||||
|
|
||||||
|
Read the project's domain glossary (`CONTEXT.md`) and any ADRs in the area you're touching first.
|
||||||
|
|
||||||
|
Then spawn a sub-agent to walk the codebase. Don't follow rigid heuristics — explore organically and note where you experience friction:
|
||||||
|
|
||||||
|
- Where does understanding one concept require bouncing between many small modules?
|
||||||
|
- Where are modules **shallow** — interface nearly as complex as the implementation?
|
||||||
|
- Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
|
||||||
|
- Where do tightly-coupled modules leak across their seams?
|
||||||
|
- Which parts of the codebase are untested, or hard to test through their current interface?
|
||||||
|
|
||||||
|
Apply the **deletion test** to anything you suspect is shallow: would deleting it concentrate complexity, or just move it? A "yes, concentrates" is the signal you want.
|
||||||
|
|
||||||
|
### 2. Present candidates as an HTML report
|
||||||
|
|
||||||
|
Write a self-contained HTML file to the OS temp directory so nothing lands in the repo. Resolve the temp dir from `$TMPDIR`, falling back to `/tmp` (or `%TEMP%` on Windows), and write to `<tmpdir>/architecture-review-<timestamp>.html` so each run gets a fresh file. Open it for the user — `xdg-open <path>` on Linux, `open <path>` on macOS, `start <path>` on Windows — and tell them the absolute path.
|
||||||
|
|
||||||
|
The report uses **Tailwind via CDN** for layout and styling, and **Mermaid via CDN** for diagrams where a graph/flow/sequence reliably communicates the structure. Mix Mermaid with hand-crafted CSS/SVG visuals — use Mermaid when relationships are graph-shaped (call graphs, dependencies, sequences), and hand-built divs/SVG when you want something more editorial (mass diagrams, cross-sections, collapse animations). Each candidate gets a **before/after visualisation**. Be visual.
|
||||||
|
|
||||||
|
For each candidate, render a card with:
|
||||||
|
|
||||||
|
- **Files** — which files/modules are involved
|
||||||
|
- **Problem** — why the current architecture is causing friction
|
||||||
|
- **Solution** — plain English description of what would change
|
||||||
|
- **Benefits** — explained in terms of locality and leverage, and how tests would improve
|
||||||
|
- **Before / After diagram** — side-by-side, custom-drawn, illustrating the shallowness and the deepening
|
||||||
|
- **Recommendation strength** — one of `Strong`, `Worth exploring`, `Speculative`, rendered as a badge
|
||||||
|
|
||||||
|
End the report with a **Top recommendation** section: which candidate you'd tackle first and why.
|
||||||
|
|
||||||
|
**Use CONTEXT.md vocabulary for the domain, and the `/codebase-design` vocabulary for the architecture.** If `CONTEXT.md` defines "Order," talk about "the Order intake module" — not "the FooBarHandler," and not "the Order service."
|
||||||
|
|
||||||
|
**ADR conflicts**: if a candidate contradicts an existing ADR, only surface it when the friction is real enough to warrant revisiting the ADR. Mark it clearly in the card (e.g. a warning callout: _"contradicts ADR-0007 — but worth reopening because…"_). Don't list every theoretical refactor an ADR forbids.
|
||||||
|
|
||||||
|
See [HTML-REPORT.md](HTML-REPORT.md) for the full HTML scaffold, diagram patterns, and styling guidance.
|
||||||
|
|
||||||
|
Do NOT propose interfaces yet. After the file is written, ask the user: "Which of these would you like to explore?"
|
||||||
|
|
||||||
|
### 3. Grilling loop
|
||||||
|
|
||||||
|
Once the user picks a candidate, call the Skill tool with "grilling" to walk the decision tree with them — constraints, dependencies, the shape of the deepened module, what sits behind the seam, what tests survive.
|
||||||
|
|
||||||
|
Side effects happen inline as decisions crystallize — call the Skill tool with "domain-modeling" to keep the domain model current as you go:
|
||||||
|
|
||||||
|
- **Naming a deepened module after a concept not in `CONTEXT.md`?** Add the term to `CONTEXT.md`. Create the file lazily if it doesn't exist.
|
||||||
|
- **Sharpening a fuzzy term during the conversation?** Update `CONTEXT.md` right there.
|
||||||
|
- **User rejects the candidate with a load-bearing reason?** Offer an ADR, framed as: _"Want me to record this as an ADR so future architecture reviews don't re-suggest it?"_ Only offer when the reason would actually be needed by a future explorer to avoid re-suggesting the same thing — skip ephemeral reasons ("not worth it right now") and self-evident ones.
|
||||||
|
- **Want to explore alternative interfaces for the deepened module?** Call the Skill tool with "codebase-design" and use its design-it-twice parallel sub-agent pattern.
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Improve Codebase Architecture"
|
||||||
|
short_description: "Find and grill architecture improvements"
|
||||||
|
policy:
|
||||||
|
allow_implicit_invocation: false
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
---
|
||||||
|
name: tdd
|
||||||
|
description: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Test-Driven Development
|
||||||
|
|
||||||
|
TDD is the red → green loop. This skill is the reference that makes that loop produce tests worth keeping: what a good test is, where tests go, the anti-patterns, and the rules of the loop. Every section applies on every cycle — consult them before and during the loop, not after.
|
||||||
|
|
||||||
|
When exploring the codebase, read `CONTEXT.md` (if it exists) so test names and interface vocabulary match the project's domain language, and respect ADRs in the area you're touching.
|
||||||
|
|
||||||
|
## What a good test is
|
||||||
|
|
||||||
|
Tests verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't. A good test reads like a specification — "user can checkout with valid cart" tells you exactly what capability exists — and survives refactors because it doesn't care about internal structure.
|
||||||
|
|
||||||
|
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
|
||||||
|
|
||||||
|
## Seams — where tests go
|
||||||
|
|
||||||
|
A **seam** is the public boundary you test at: the interface where you observe behavior without reaching inside. Tests live at seams, never against internals.
|
||||||
|
|
||||||
|
**Test only at pre-agreed seams.** Before writing any test, write down the seams under test and confirm them with the user. No test is written at an unconfirmed seam. You can't test everything — agreeing the seams up front is how testing effort lands on the critical paths and complex logic instead of every edge case.
|
||||||
|
|
||||||
|
Ask: "What's the public interface, and which seams should we test?"
|
||||||
|
|
||||||
|
When the shape of that interface is itself in question — how deep the module is, where the seam belongs, what the interface should expose — call the Skill tool with "codebase-design" for the vocabulary. It is the shared source of the module, interface, depth, seam, adapter, leverage and locality terms, and it is a reference to consult, not a session to run.
|
||||||
|
|
||||||
|
## Anti-patterns
|
||||||
|
|
||||||
|
- **Implementation-coupled** — mocks internal collaborators, tests private methods, or verifies through a side channel (querying the database instead of using the interface). The tell: the test breaks when you refactor but behavior hasn't changed.
|
||||||
|
- **Tautological** — the assertion recomputes the expected value the way the code does (`expect(add(a, b)).toBe(a + b)`, a snapshot derived by hand the same way, a constant asserted equal to itself), so it passes by construction and can never disagree with the code. Expected values must come from an independent source of truth — a known-good literal, a worked example, the spec.
|
||||||
|
- **Horizontal slicing** — writing all tests first, then all implementation. Bulk tests verify _imagined_ behavior: you test the _shape_ of things rather than user-facing behavior, the tests go insensitive to real changes, and you commit to test structure before understanding the implementation. Work in **vertical slices** instead — one test → one implementation → repeat, each test a **tracer bullet** that responds to what the last cycle taught you.
|
||||||
|
|
||||||
|
## Rules of the loop
|
||||||
|
|
||||||
|
- **Red before green.** Write the failing test first, then only enough code to pass it. Don't anticipate future tests or add speculative features.
|
||||||
|
- **One slice at a time.** One seam, one test, one minimal implementation per cycle.
|
||||||
|
- **Refactoring is not part of the loop.** It belongs to the review stage (see the `code-review` skill), not the red → green implementation cycle.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "TDD"
|
||||||
|
short_description: "Test-driven red-green-refactor"
|
||||||
@@ -0,0 +1,59 @@
|
|||||||
|
# When to Mock
|
||||||
|
|
||||||
|
Mock at **system boundaries** only:
|
||||||
|
|
||||||
|
- External APIs (payment, email, etc.)
|
||||||
|
- Databases (sometimes - prefer test DB)
|
||||||
|
- Time/randomness
|
||||||
|
- File system (sometimes)
|
||||||
|
|
||||||
|
Don't mock:
|
||||||
|
|
||||||
|
- Your own classes/modules
|
||||||
|
- Internal collaborators
|
||||||
|
- Anything you control
|
||||||
|
|
||||||
|
## Designing for Mockability
|
||||||
|
|
||||||
|
At system boundaries, design interfaces that are easy to mock:
|
||||||
|
|
||||||
|
**1. Use dependency injection**
|
||||||
|
|
||||||
|
Pass external dependencies in rather than creating them internally:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Easy to mock
|
||||||
|
function processPayment(order, paymentClient) {
|
||||||
|
return paymentClient.charge(order.total);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Hard to mock
|
||||||
|
function processPayment(order) {
|
||||||
|
const client = new StripeClient(process.env.STRIPE_KEY);
|
||||||
|
return client.charge(order.total);
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**2. Prefer SDK-style interfaces over generic fetchers**
|
||||||
|
|
||||||
|
Create specific functions for each external operation instead of one generic function with conditional logic:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// GOOD: Each function is independently mockable
|
||||||
|
const api = {
|
||||||
|
getUser: (id) => fetch(`/users/${id}`),
|
||||||
|
getOrders: (userId) => fetch(`/users/${userId}/orders`),
|
||||||
|
createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
|
||||||
|
};
|
||||||
|
|
||||||
|
// BAD: Mocking requires conditional logic inside the mock
|
||||||
|
const api = {
|
||||||
|
fetch: (endpoint, options) => fetch(endpoint, options),
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
The SDK approach means:
|
||||||
|
- Each mock returns one specific shape
|
||||||
|
- No conditional logic in test setup
|
||||||
|
- Easier to see which endpoints a test exercises
|
||||||
|
- Type safety per endpoint
|
||||||
@@ -0,0 +1,77 @@
|
|||||||
|
# Good and Bad Tests
|
||||||
|
|
||||||
|
## Good Tests
|
||||||
|
|
||||||
|
**Integration-style**: Test through real interfaces, not mocks of internal parts.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// GOOD: Tests observable behavior
|
||||||
|
test("user can checkout with valid cart", async () => {
|
||||||
|
const cart = createCart();
|
||||||
|
cart.add(product);
|
||||||
|
const result = await checkout(cart, paymentMethod);
|
||||||
|
expect(result.status).toBe("confirmed");
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
Characteristics:
|
||||||
|
|
||||||
|
- Tests behavior users/callers care about
|
||||||
|
- Uses public API only
|
||||||
|
- Survives internal refactors
|
||||||
|
- Describes WHAT, not HOW
|
||||||
|
- One logical assertion per test
|
||||||
|
|
||||||
|
## Bad Tests
|
||||||
|
|
||||||
|
**Implementation-detail tests**: Coupled to internal structure.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// BAD: Tests implementation details
|
||||||
|
test("checkout calls paymentService.process", async () => {
|
||||||
|
const mockPayment = jest.mock(paymentService);
|
||||||
|
await checkout(cart, payment);
|
||||||
|
expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
Red flags:
|
||||||
|
|
||||||
|
- Mocking internal collaborators
|
||||||
|
- Testing private methods
|
||||||
|
- Asserting on call counts/order
|
||||||
|
- Test breaks when refactoring without behavior change
|
||||||
|
- Test name describes HOW not WHAT
|
||||||
|
- Verifying through external means instead of interface
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// BAD: Bypasses interface to verify
|
||||||
|
test("createUser saves to database", async () => {
|
||||||
|
await createUser({ name: "Alice" });
|
||||||
|
const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
|
||||||
|
expect(row).toBeDefined();
|
||||||
|
});
|
||||||
|
|
||||||
|
// GOOD: Verifies through interface
|
||||||
|
test("createUser makes user retrievable", async () => {
|
||||||
|
const user = await createUser({ name: "Alice" });
|
||||||
|
const retrieved = await getUser(user.id);
|
||||||
|
expect(retrieved.name).toBe("Alice");
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
**Tautological tests**: Expected value restates the implementation, so the test passes by construction.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// BAD: Expected value is recomputed the way the code computes it
|
||||||
|
test("calculateTotal sums line items", () => {
|
||||||
|
const items = [{ price: 10 }, { price: 5 }];
|
||||||
|
const expected = items.reduce((sum, i) => sum + i.price, 0);
|
||||||
|
expect(calculateTotal(items)).toBe(expected);
|
||||||
|
});
|
||||||
|
|
||||||
|
// GOOD: Expected value is an independent, known literal
|
||||||
|
test("calculateTotal sums line items", () => {
|
||||||
|
expect(calculateTotal([{ price: 10 }, { price: 5 }])).toBe(15);
|
||||||
|
});
|
||||||
|
```
|
||||||
@@ -7,12 +7,11 @@ tool_output_token_limit = 25000
|
|||||||
# With tool_output_token_limit=25000 ⇒ 273000 - (25000 + 15000) = 233000
|
# With tool_output_token_limit=25000 ⇒ 273000 - (25000 + 15000) = 233000
|
||||||
model_auto_compact_token_limit = 233000
|
model_auto_compact_token_limit = 233000
|
||||||
suppress_unstable_features_warning = true
|
suppress_unstable_features_warning = true
|
||||||
notify = ["/home/sudacode/.codex/scripts/codex-notify"]
|
|
||||||
sandbox_mode = "workspace-write"
|
sandbox_mode = "workspace-write"
|
||||||
service_tier = "default"
|
service_tier = "default"
|
||||||
|
|
||||||
[tui]
|
[tui]
|
||||||
notifications = ["agent-turn-complete"]
|
notifications = ["agent-turn-complete", "approval-requested"]
|
||||||
notification_condition = "always"
|
notification_condition = "always"
|
||||||
|
|
||||||
[tui.model_availability_nux]
|
[tui.model_availability_nux]
|
||||||
@@ -247,9 +246,6 @@ enabled = true
|
|||||||
[plugins."chrome@openai-bundled"]
|
[plugins."chrome@openai-bundled"]
|
||||||
enabled = true
|
enabled = true
|
||||||
|
|
||||||
[plugins."subminer-workflow@subminer-local"]
|
|
||||||
enabled = true
|
|
||||||
|
|
||||||
[plugins."coderabbit@openai-curated"]
|
[plugins."coderabbit@openai-curated"]
|
||||||
enabled = true
|
enabled = true
|
||||||
|
|
||||||
|
|||||||
@@ -234,8 +234,7 @@
|
|||||||
// "20": "十",
|
// "20": "十",
|
||||||
},
|
},
|
||||||
"persistent-workspaces": {
|
"persistent-workspaces": {
|
||||||
"*": 10,
|
"*": [1, 2, 3, 4, 5, 6, 7, 8, 9, 10],
|
||||||
// "*": 5,
|
|
||||||
},
|
},
|
||||||
"sort-by": "number",
|
"sort-by": "number",
|
||||||
"all-outputs": false,
|
"all-outputs": false,
|
||||||
|
|||||||
Reference in New Issue
Block a user