2,020 messages across 93 sessions (115 total) | 2026-08-05 to 2026-09-29
At a Glance
What's working: You've built a disciplined verify-then-fix loop for MR reviews. Claude checks each finding against the code before touching it, pushes with tests, and resolves threads with explanations. That loop has caught reviewer mistakes and left many MRs ready to merge. You also treat artifacts as team infrastructure: explainers for non-engineers, meeting reports, colleague guides, and handoff docs you dry-run so a fresh session can pick up long efforts. You audit your own tooling too, including /dual-review itself and the causes behind rising review findings, which keeps the process getting sharper. Impressive Things You Did →
What's hindering you: On Claude's side, the biggest issue is stating unverified things as fact: a live config called dead, tests said to exist, inflated file counts, and false dependency claims. It also reached for subagents and elaborate doctrine when you wanted something lighter. On your side, shared resources keep colliding: edits landing in a repo someone else was using, Chrome DevTools profiles locked by another session, and /resume picking sessions still running in the background. Harness and meta-work also tended to start without a scoped plan, which led to one expensive wasted run. Where Things Go Wrong →
Quick wins to try: Add a Hook that blocks Edit/Write outside a worktree and gives each session its own Chrome profile. That removes a whole class of collisions without you having to remember. Since you already live in skills, fold your review-feedback pipeline (verify, fix, test, push, reply, resolve) into one lean custom skill with a built-in 'no subagents unless asked' default. Try Headless Mode for recurring jobs like weekly MR status artifacts or dual-review retries after Codex quota resets. Features to Try →
Ambitious workflows: Your review, fix, and resolve pipeline is well suited to running fully autonomously per MR. Given just an MR number, it would gate, review, verify, fix, wait for green CI, and resolve threads. You'd only see a single decision artifact for spec questions it can't settle. Beyond that, prepare for a parallel issue-to-MR fleet where each agent gets its own worktree, branch, and browser profile and iterates against Spectra specs until green. Pair it with a claims-verifier that backs every statement in an MR description or artifact with a runnable probe before anything reaches a colleague. On the Horizon →
This was the dominant workflow. The user ran /dual-review and pr-review-toolkit reviews across many GitLab MRs (!540–!588), posted the findings, and then had Claude verify, fix and push the review items before replying to and resolving threads. Claude used parallel review agents, worked around Codex quota exhaustion, resolved merge conflicts, and checked fixes on preview and staging.
Skyline/Tequila Feature Development & Bug Fixes~10 sessions
Claude built product features and fixes for the skyline and tequila apps. Examples include re-rejection support with rejection history, a UPC lookup bug fix done with TDD, and restoring filter persistence across tabs. Other work covered pitch-package export failures (a 2.4GB WAV RangeError diagnosed with live Chrome DevTools probes), a GTM/GA4 A/B test, and resyncing MR 492 to a new contract-review spec. Changes typically went through Spectra specs, tests, worktrees and draft MRs.
Claude Code Harness, Skills & Config Tooling~10 sessions
The user tuned their Claude Code setup: improving the /dual-review and issue-to-mr skills, building a skill router and token-saving experiments, adding repo rules such as 'avoid subagents' and a comment policy, and cleaning up memory from 131 entries to 8. This area had the most friction. Problems included over-engineered doctrine, miscommunication that wasted an expensive run, and heavy subagent use the user found excessive.
Claude upgraded the repo to Spectra 3.0, fixed 28 invalid specs, and investigated why Spectra couldn't see specs in git worktrees, which led to a feature-request issue. It also explained worktree placement in artifacts and implemented worktree lane isolation. One MR wrongly declared openspec/config.yaml dead until a user-driven probe proved otherwise.
Status Reports, Explainers & Dev Environment~9 sessions
Claude produced HTML artifacts for meetings and non-engineers: weekly MR status reports, release feature surveys for demos, feature explainers, and an analysis showing that rising review findings came from new AI review bots. It also ran the 1.6.6 release, where it caught unpushed commits before a worktree delete. On the environment side, it built a split-tunnel VPN setup with dnsmasq and diagnosed a macOS input-method issue.
What You Wanted
Address Code Review Feedback
8
Post Review To Mr
7
Conceptual Question
7
Code Review
7
Update Artifact
6
Check Interrupted Tasks
4
Top Tools Used
Bash
7732
Agent
453
Edit
445
Mcp Chrome-Devtools Evaluate Script
281
Write
208
Read
193
Languages
HTML
297
TypeScript
295
Markdown
259
JavaScript
43
Python
17
Shell
13
Session Types
Multi Task
34
Iterative Refinement
7
Single Task
5
Quick Question
2
Exploration
2
How You Use Claude Code
You work with Claude Code as a delegating engineering lead running a review-and-fix pipeline, not as a pair programmer typing alongside it. Most of your sessions start with a compact, outcome-shaped directive, such as "verify and fix the review findings on !552," "dual-review these three MRs and post," or "implement both batches from handoff §3.0 without touching the handoff doc." Then you let Claude run long stretches: tests, rebases, CI, thread resolution, artifacts. The numbers show this: 7,732 Bash calls, 453 Agent dispatches and 230 commits across 93 sessions. You chain work across sessions with handoff docs, artifacts and Spectra changes as durable memory. MR 492, for example, carried across several sessions through a handoff doc listing pending questions for Chris. You also ask for artifacts to be self-sufficient for a new session, as you did with the comment-policy inventory, which was checked with two dry runs.
You tend to intervene at checkpoints rather than micromanage each step, and when you do, it's a sharp course correction grounded in operational judgment. You redirected work into a worktree when Claude edited the main repo while someone else was using it. You pointed out that rule edits overlapped the existing !588. You declined an unnecessary tsc fix. You told Claude to stop using subagents when the multi-agent review process felt like overkill, and then had that turned into a repo rule in !591. Your own probe test caught Claude wrongly calling openspec/config.yaml dead. You also push back on verbosity and overblown concerns. The friction record points one way: buggy code (31) and wrong approach (29) dominate. You tolerate self-corrected slips well, but you get openly frustrated when Claude over-engineers, makes unverified claims, or wastes your resources. The harness/Jev session is the clearest case: a false dependency claim and a narrowed scope cost you an expensive /dual-review run. Your anger over the unused 7/5 doctrine is another.
You also invest heavily in meta-work on your own tooling. You audit prompts and rules (117 findings), evaluate /dual-review with evidence, prune memory from 131 entries to 8, strip the global CLAUDE.md, and ask why review findings are rising. You treat Claude Code as a system you actively tune for cost and reliability. When that tuning pays off, you say so warmly: "比我過年大掃除還舒爽" (roughly "more refreshing than my New Year's deep clean"). Outside that loop, you also send quick conceptual questions and write non-engineer explainers for colleagues and meetings, so Claude doubles as your communication layer to the team.
Key pattern: You hand Claude terse, MR-centric directives, let it run long verify-fix-push-resolve cycles across worktrees and handoff docs, and step in sharply when it over-engineers, makes unverified claims, or wastes your tokens.
User Response Time Distribution
2-10s
58
10-30s
153
30s-1m
125
1-2m
207
2-5m
280
5-15m
311
>15m
212
Median: 182.6s • Average: 470.4s
Multi-Clauding (Parallel Sessions)
99
Overlap Events
80
Sessions Involved
32%
Of Messages
You run multiple Claude Code sessions simultaneously. Multi-clauding is detected when sessions
overlap in time, suggesting parallel workflows.
User Messages by Time of Day
Morning (6-12)
865
Afternoon (12-18)
713
Evening (18-24)
205
Night (0-6)
237
Tool Errors Encountered
Command Failed
130
Other
115
User Rejected
12
Edit Failed
4
File Not Found
2
Impressive Things You Did
Across 93 sessions over two months, you've built a review-driven, MR-centric workflow on your tequila/skyline repos, with 230 commits and most sessions fully or mostly achieving their goals.
Verify-then-fix code review loops
You've turned MR review feedback into a disciplined pipeline. You have Claude independently verify each finding before fixing it, then push with tests and resolve the threads with explanations. This caught reviewer mistakes (such as the wrong numbers on !540), exposed issues the reviews missed, and repeatedly left MRs like !547, !551 and !565 ready to merge, which makes it one of your most consistently 'essential' patterns.
Resilient multi-model dual reviews
You run /dual-review across batches of MRs with gating logic that skips MRs not worth reviewing. When Codex quota runs out, you fall back to Claude reviewers or schedule auto-retries, and you approve findings before anything gets posted. You also audit the tooling itself: you asked for evidence-based evaluations of /dual-review and investigated why review findings were rising, so the review process keeps getting better.
Artifacts as durable team knowledge
You regularly turn technical work into polished artifacts for different audiences: explainers for non-engineers, weekly MR status reports for meetings, VPN split-tunnel guides for colleagues, and handoff docs that are dry-run tested so a fresh session can pick them up. Combined with Spectra specs and handoff docs, this lets you spread long efforts like MR 492 across sessions without losing context. It also makes Claude's output useful to your wider team, not just to you.
What Helped Most (Claude's Capabilities)
Multi-file Changes
17
Good Debugging
14
Proactive Help
7
Good Explanations
7
Fast/Accurate Search
2
Correct Code Edits
2
Outcomes
Not Achieved
3
Partially Achieved
3
Mostly Achieved
14
Fully Achieved
30
Where Things Go Wrong
Your friction clusters around Claude making confident claims it hadn't fully verified, defaulting to heavyweight multi-agent processes you didn't want, and working in shared or contended environments (main repo, Chrome profile, resumed sessions) without checking first.
Unverified claims stated as fact
Claude often asserted things about configs, tests, or dependencies before checking every path. You then had to catch the mistakes through probes or reviews. You could require Claude to show evidence, such as a file path, a command's output, or the GUI/CLI surface it tested, before it calls something dead, existing, or required.
Claude declared openspec/config.yaml dead after testing only the CLI and skills, not the GUI app. You deleted it, and it had to be restored after your probe test and a reviewer proved it was live.
In the Jev/harness session, Claude claimed hooks depended on !571 and let you run an expensive /dual-review believing Jev would save tokens when it wasn't wired in. That cost you a wasted run and caused repeated frustration.
Over-engineered multi-agent workflows
Claude reached for subagents, dual reviews and elaborate doctrines more than you wanted. Network stalls then compounded the overhead and made the heaviness worse. The repo-level 'do it yourself, avoid subagents' rule you added was the right move. Consider stating the scope ('quick fix, no reviewers') at the start of each task.
You had to ask why Claude kept dispatching review agents and ended up opening MR !591 to ban subagent-heavy reviews. Before that, stalled agents were misdiagnosed as a 'bloated context' problem.
The 7/5 doctrine in your global Claude config was over-engineered and never took effect. You were angry, and it had to be fully removed along with the global CLAUDE.md contents and the restic backup.
Shared-environment and session collisions
Work got blocked or risky because Claude didn't check who else was using the repo, the browser profile, or the session. Tell Claude to always start in a worktree, and close or dedicate Chrome DevTools profiles per session. That avoids these repeated stalls.
Claude edited files directly in the main repo while someone else was using it, so you had to ask for the work to be moved into a worktree mid-task. The same thing happened again during the /dual-review evaluation.
Another Claude session held the Chrome DevTools profile, which blocked staging verification on MR 492 and the GTM/GA4 console setup. Separately, two /resume attempts failed because session 06b73c90 was still running in the background.
Primary Friction Types
Buggy Code
31
Wrong Approach
29
Misunderstood Request
13
External Interruption
7
Tool Environment Issues
4
Tool Environment Issue
3
Inferred Satisfaction (model-estimated)
Frustrated
4
Dissatisfied
19
Likely Satisfied
167
Satisfied
13
Happy
2
Unsure
1
Existing CC Features to Try
Suggested CLAUDE.md Additions
Just copy this into Claude Code to add it to your CLAUDE.md.
In several sessions (UPC fix, tab-filter change, dual-review evaluation) the user had to ask for work to be moved into a worktree, and one cleanup nearly deleted 8 unpushed commits.
The user asked why Claude kept dispatching review agents, told it to stop using the code-reviewer, and called the multi-agent process overkill, even though the Agent tool was used 453 times.
Wrong claims caused real rework: openspec/config.yaml called dead, nonexistent tab-switch tests, false hook dependencies on !571, 36 files vs the actual 5, 14 vs 15 suggestions. The user also pushed back on verbosity and overblown concerns.
zsh quoting bugs, locked Chrome profiles and Codex quota failures each came up in 3 or more sessions and cost time every time.
Just copy this into Claude Code and it'll set it up for you.
Hooks
Shell commands that run automatically before or after tool calls.
Why for you: A PreToolUse hook can block edits in the main checkout, so you no longer have to say 'move this to a worktree'. This hook blocks when your current directory is the main repo root, so start sessions for worktree tasks from inside the worktree.
// .claude/settings.json
{
"hooks": {
"PreToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "if [ \"$(git rev-parse --git-dir)\" = \".git\" ]; then echo 'Blocked: create a worktree before editing the main checkout' >&2; exit 2; fi"
}
]
}
]
}
}
Custom Skills
Turn your most common multi-step workflow into a single /command.
Why for you: Your top goal is address_code_review_feedback, and the steps are the same every time: verify each finding, fix, test, push, update the MR description, reply to and resolve threads. A /fix-review skill would make that consistent, and it can build in the 'no extra subagents' and 'verify first' rules.
# .claude/skills/fix-review/SKILL.md
---
name: fix-review
description: Verify and fix review findings on a GitLab MR
---
1. `glab mr view $ARGUMENTS --comments` and list every unresolved finding.
2. Verify each finding against the code yourself (no subagents). Mark each confirmed, rejected or partial, with evidence.
3. Enter or create the MR's worktree. Fix the confirmed findings with tests.
4. Run the test suite and lint. Push.
5. Update the MR description if scope changed.
6. Reply to each thread with what was done or why it was rejected, then resolve it. Report a concise summary.
Headless Mode
Run Claude non-interactively from scripts, cron or CI.
Why for you: Dual reviews and weekly MR status reports come up often, and they keep hitting Codex quota limits or blocking your interactive sessions. Running them headless in the background or in GitLab CI frees your main session and makes quota retries automatic.
# review every open MR assigned to you, one at a time
for mr in $(glab mr list --reviewer=@me -F json | jq -r '.[].iid'); do
claude -p "/dual-review !$mr — do not post to the MR; save the report to reviews/$mr.md" \
--allowedTools "Bash,Read,Skill" > "logs/review-$mr.log" 2>&1
done
New Ways to Use Claude Code
Just copy this into Claude Code and it'll walk you through it.
Avoid Chrome profile collisions
Give each session its own browser profile for Chrome DevTools MCP so parallel sessions stop blocking each other.
At least 4 sessions (MR 492 staging verification, the tab-filter walkthrough, the GTM/GA4 setup, the weekly report render) were blocked because another session held the Chrome DevTools profile. You run many sessions in parallel, so a shared profile is a recurring bottleneck. The chrome-devtools MCP server supports an isolated, temporary profile per launch. The catch is that isolated profiles don't keep staging logins, so keep one persistent profile for authenticated work.
Paste into Claude Code:
Reconfigure my chrome-devtools MCP server so each Claude session uses an isolated profile (`claude mcp add chrome-devtools -- npx chrome-devtools-mcp@latest --isolated`). Also keep a second named server with a persistent profile for staging-login work, and explain when to use each.
Stop resuming sessions that are still running
Two sessions failed outright because /resume picked a session that was still running in the background.
Both 'not_achieved' resume sessions had the same cause: session 06b73c90 was still running, so no context loaded. When you work across many parallel sessions with handoff docs, it's easy to lose track of what's still running. Check or attach to the running session first, or rely on the handoff artifacts you already write so a fresh session can continue without /resume.
Paste into Claude Code:
Read the latest handoff doc for this MR and continue from its 'Next steps' section. Before doing anything, list which tasks are done vs pending according to the doc and the git log, and confirm with me.
Start ambitious meta-work with a scoped plan
For harness, tooling and config projects, have Claude read the relevant docs and confirm scope and dependencies before running expensive experiments.
The Jev harness session caused the most frustration of the period. The scope quietly narrowed to skill routing, false dependency claims led you into a wasted /dual-review run, and the effort/model routing you actually wanted was evaluated late. The 7/5 doctrine had a similar problem: over-engineered and never took effect. A short written plan listing goals, prerequisites and what is actually wired in yet would have caught these problems early. Your successful prompt-audit session worked because findings were delivered first and you chose the fixes.
Paste into Claude Code:
Before implementing anything: read the official Claude Code docs relevant to this goal and write a 1-page plan. Include (1) my goal in your words, (2) what is already wired in vs not, with evidence, (3) prerequisites and true dependencies, (4) the cheapest experiment that proves value. Wait for my approval before running anything that costs tokens.
On the Horizon
Your work has already moved from Claude as a pair programmer to Claude as an MR pipeline operator: dual-reviews, fix-the-findings loops, stacked MRs and artifacts. The next step is closing those loops on their own, with tests and live probes as the judge instead of you.
Autonomous review-to-merge-ready loop per MR
Your most common goals are code review, posting reviews and fixing review feedback, and they already form one repeatable pipeline. A single headless run could take an MR number and go straight through: gate, review, verify each finding against the code, fix with mutation-tested coverage, push, wait for green CI, and resolve threads with evidence. You would only be pulled in for spec decisions it cannot settle, arriving as one decision artifact rather than a string of back-and-forth messages.
Getting started: Wrap your existing /dual-review and glab skills in a single skill that ends at a hard stopping condition: CI green, every thread resolved or escalated. Run it with `claude -p` in a dedicated worktree so it never touches the shared main checkout.
Paste into Claude Code:
Run the full review-to-merge-ready loop for MR !<N> in a fresh git worktree (never edit the main checkout). Steps:
1. Gate: decide whether the MR is worth a full review and explain why.
2. Review: run /dual-review. If Codex hits its quota, fall back to Claude reviewers and don't wait.
3. Verify every finding against the actual code and spec. Classify each as confirmed, rejected (with evidence) or needs-human-decision. Never assert something is dead or unused unless you have probed every consumer: CLI, GUI and skills.
4. Fix confirmed findings test-first. Run the full test suite and a mutation check on the new tests.
5. Push, then poll CI until it is green. Fix any failures yourself.
6. Reply to each thread with the commit SHA and evidence, then resolve it.
Stop only when CI is green and every thread is resolved or escalated. Then publish one short artifact listing only the needs-human-decision items, each with options and your recommendation. Keep the prose terse.
Parallel issue-to-MR fleet with isolated lanes
You already run issue-to-mr, stacked MRs and worktree lane isolation. The next step is a fleet: an orchestrator splits your open, unblocked issues across parallel agents, each in its own worktree, branch and Chrome profile. Each lane iterates against the tests and the Spectra specs until green, then opens a draft MR. Blocked issues come back as a triage table instead of wasted runs, and the Chrome profile lock stops blocking verification.
Getting started: Use the Agent tool from a coordinator session, with one git worktree and one `--user-data-dir` chrome-devtools profile per lane. Cap concurrency at 3 or 4, and have each lane write a heartbeat file so network-stalled agents can be detected and resumed.
Paste into Claude Code:
Act as an orchestrator for our open issues in tequila and skyline.
1. List the open issues and classify each as ready, blocked (name the dependency, e.g. backend) or needs-clarification.
2. For up to 4 ready issues, spawn one subagent each. Each gets its own git worktree off staging, its own branch and its own isolated Chrome profile directory.
3. Each lane must: write or update the Spectra change, implement test-first, iterate until the full suite passes, verify the UI on staging or preview with chrome-devtools in its own profile, then open a draft MR that links the issue.
4. Each lane writes status to .lanes/<issue>.json every step. If a lane has no update for 10 minutes, resume or restart it.
5. Before opening MRs, check for overlap with existing open MRs and stack or rebase as needed.
Finish with one artifact: a table of issue, lane result, MR link, test count, and what needs my input.
Self-verifying claims before anything gets published
Your biggest friction sources were buggy code, wrong approaches and confident false claims: a live config called 'dead', '36 files' that was really 5, a wrong sync date, and tests described as existing when they didn't. An adversarial verifier agent could extract every factual claim from a draft MR description, artifact or review. It would then check each one with a runnable probe (grep, test run, live query, git log) before anything reaches you or a colleague. Artifacts and MR descriptions would ship with a claims ledger in which every claim has evidence.
Getting started: Make it a skill or a Stop hook that runs on any artifact or MR-description draft. It spawns a fresh-context verifier subagent, which avoids the same-context anchoring you already identified, and blocks publishing if any claim is unverified.
Paste into Claude Code:
Before publishing this artifact or MR description, run a claims audit.
1. Extract every factual claim: counts, dates, 'X is unused/dead', 'tests already cover Y', 'depends on !NNN', and behavior descriptions.
2. For each claim, spawn a fresh-context verifier that has NOT seen your reasoning. Give it only the claim and the repo. It must design and run a concrete probe (grep excluding type-only imports, git log, running the test, checking every consumer including GUI apps, a live staging query) and return verified, false or unverifiable with the command output.
3. Fix or remove false claims. Soften unverifiable ones and say so explicitly.
4. Append a collapsed 'Evidence' section mapping each claim to its probe.
Only publish when zero claims are false. Report to me in 3 lines or fewer: how many claims were checked, how many were corrected, and anything unverifiable.
"After Claude cut the user's bloated auto-memory from 131 entries down to 8, the user declared it "比我過年大掃除還舒爽", which roughly means "more refreshing than my New Year's deep clean.""
This came from a session to disable auto-memory and do a thorough cleanup. Claude moved the useful knowledge into the official glab skill, merged the MRs and kept a visual artifact updated throughout. It was one of only two 'happy' reactions in the whole dataset.