Newer Claude CLI versions read from stdin even in -p (print) mode.
This causes hangs in both execution modes:
- Background mode: OS sends SIGTTIN to suspend the process; the
progress-monitoring kill -0 loop runs forever since it returns
true for stopped (not just running) processes.
- Live/streaming mode (--monitor/--live): the foreground process
blocks waiting for terminal input it never receives.
Fix: add < /dev/null to both the modern CLI background execution path
and the live streaming pipeline, consistent with legacy mode which
already redirects stdin from $PROMPT_FILE.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Disable live mode when build_claude_command fails (not just empty array)
- Strengthen safety check to verify use_modern_cli=true
- Fix case-sensitive sed range in test (execute→Execute)
- Update CLAUDE.md: auto-switch docs, test counts (484→490)
Live/monitor mode crashed with `stdbuf: unrecognized option '--verbose'`
when CLAUDE_OUTPUT_FORMAT="text" because build_claude_command() was only
called inside a JSON-only gate, leaving CLAUDE_CMD_ARGS empty.
Three coordinated fixes:
- Override text→json when live mode is active (stream-json requires JSON)
- Always call build_claude_command() regardless of output format
- Add safety check for empty CLAUDE_CMD_ARGS before live mode construction
Also updates help text for --live and --output-format to document the
auto-switch behavior.
The OPEN state was terminal — once triggered, it persisted across
restarts with no automatic recovery. This adds two recovery mechanisms:
1. Cooldown timer (default): OPEN → HALF_OPEN after CB_COOLDOWN_MINUTES
(default 30). The existing HALF_OPEN logic handles recovery or re-trip.
2. Auto-reset option: CB_AUTO_RESET=true bypasses cooldown, resets to
CLOSED on startup for fully unattended operation.
Changes:
- Add parse_iso_to_epoch() to lib/date_utils.sh (cross-platform)
- Add cooldown + auto-reset logic to init_circuit_breaker()
- Add opened_at field to state file when entering/staying OPEN
- Add --auto-reset-circuit CLI flag and .ralphrc config vars
- Add 19 tests in test_circuit_breaker_recovery.bats
- Update CLAUDE.md and README.md documentation
* fix: progress detection improvements (#141, #144)
- Fix checkbox regex to exclude date entries like [2026-01-29] (#144)
- Add git commit detection: files changed in commits now count as progress (#141)
- Add 13 regression tests for progress detection and checkbox regex
- Update test count from 452 to 465
Fixes#141, Fixes#144
* fix(test): increase grep context to capture echo -1 line
* fix: count both committed and working tree changes as progress
When commits are made, now unions:
- Files changed in commits (loop_start_sha..current_sha)
- Unstaged changes (git diff HEAD)
- Staged changes (git diff --cached)
Uses sort -u to deduplicate before counting.
---------
Co-authored-by: Test User <test@example.com>
The permission_denials array from Claude CLI has commands nested under
tool_input.command, not directly as .command:
Before: jq '[.permission_denials[].command]'
After: jq '[.permission_denials[].tool_input.command]'
This fixes the "Permission denied for N command(s)" message to actually
display the denied commands instead of showing empty/unknown.
Also updated test fixture to match real Claude CLI output structure.
When Claude Code is denied permission to execute commands (e.g., npm install),
Ralph now detects this from the permission_denials array in the JSON output
and halts the loop immediately with clear guidance for the user.
Changes:
- Add permission denial detection to parse_json_response() in response_analyzer.sh
- Extract permission_denials array from Claude Code JSON output
- Track has_permission_denials, permission_denial_count, denied_commands
- Add analyze_response() support for permission denial fields
- Add permission denial exit condition to should_exit_gracefully() in ralph_loop.sh
- Permission denial takes highest priority among exit conditions
- Display helpful guidance for updating ALLOWED_TOOLS in .ralphrc
- Update circuit breaker with CB_PERMISSION_DENIAL_THRESHOLD=2
- Track consecutive_permission_denials in state file
- Open circuit after 2 consecutive loops with permission denials
- Add 11 new TDD tests (6 in test_json_parsing.bats, 5 in test_exit_detection.bats)
- Update documentation in CLAUDE.md and README.md
Test count: 452 (up from 452 - added 11 new tests)
Fixes#101
Co-authored-by: Test User <test@example.com>
* fix(setup): create .ralphrc with consistent tool permissions (#136)
- Update default ALLOWED_TOOLS in ralph_loop.sh to include Edit,
Bash(npm *), and Bash(pytest) for test execution capability
- Make setup.sh generate .ralphrc file using same permissions as
ralph-enable, ensuring consistency between initialization paths
- Add 8 new TDD tests for .ralphrc creation and ALLOWED_TOOLS defaults
- Update documentation in README.md and CLAUDE.md
This fixes the mismatch where PROMPT.md instructs the model to run
tests, but the default permissions didn't allow it. Now both
ralph-setup and ralph-enable create projects with identical tool
permissions.
Test count: 440 (up from 424)
* fix: address PR review feedback
- Update version badges from v0.10.1 to v0.11.2 (README.md)
- Update test count badges from 310 to 440 (README.md, CLAUDE.md)
- Fix .ralphrc generator label: use sed to replace "ralph enable"
with "ralph-setup" when using generate_ralphrc() from library
- Add v0.11.2 changelog entry to CLAUDE.md
* docs(readme): comprehensive update for v0.11.2
- Reorganize Recent Improvements with v0.11.x versions prominent
- Add ralph-enable wizard section with full documentation
- Add .ralphrc configuration section with example
- Update Quick Start to show ralph-enable as Option A (recommended)
- Update test counts to 440 across 15 files
- Collapse v0.9.x versions into expandable details section
- Add new features to What's Working Now section
- Link to issue #138 for automated badge updates
- Update Command Reference with new commands
---------
Co-authored-by: Test User <test@example.com>
When using --monitor flag, the tmux session now correctly forwards all
CLI parameters to the inner ralph_loop.sh execution instead of only
--calls and --prompt.
Added forwarding for:
- --output-format (json/text)
- --verbose
- --timeout
- --allowed-tools
- --no-continue
- --session-expiry
Added 8 new tests for parameter forwarding validation.
Test count: 329 (up from 321)
Fixes#120
Co-Authored-By: Claude <noreply@anthropic.com>
Move variable exports before first use in the #91 regression test to
prevent "No such file or directory" failures in CI environments where
setup() state may not persist.
Co-authored-by: Test User <test@example.com>
Replace confidence-based heuristic in update_exit_signals() with explicit
EXIT_SIGNAL checking. JSON mode always has confidence >= 70 due to
deterministic scoring, causing completion_indicators to fill after 5 loops
and triggering premature exits even when Claude sets EXIT_SIGNAL: false.
Changes:
- lib/response_analyzer.sh: Check exit_signal == "true" instead of
confidence >= 60 when updating completion_indicators array
- ralph_loop.sh: Update safety circuit breaker comment to reflect that
completion_indicators now only accumulates on EXIT_SIGNAL=true
- tests/unit/test_exit_detection.bats: Add 4 TDD tests (32-35) validating
the fix for update_exit_signals() behavior
- CLAUDE.md: Document fix as v0.11.1, update test counts (420 → 424)
Test count: 424 passing (100% pass rate)
Co-authored-by: Test User <test@example.com>
* refactor(naming): remove @ prefix from on-disk filenames
BREAKING CHANGE: Renames @fix_plan.md → fix_plan.md and @AGENT.md → AGENT.md
This change improves POSIX compliance and compatibility with command-line tools.
The @ prefix was originally used to avoid naming conflicts, but with the .ralph/
folder structure introduced in v0.10.0, this convention is no longer necessary.
Changes:
- Update all scripts to use new naming (fix_plan.md, AGENT.md)
- Update migration script to handle both old and new naming conventions
- Update templates to use new naming
- Update all documentation references
- Update all 420 tests to expect new naming (TDD approach)
Migration:
- Existing projects with @-prefixed files will be automatically renamed
when running ralph-migrate
- Projects already using the new naming will continue to work unchanged
* docs(claude): update file naming conventions section
* fix(migrate): prevent orphaned @-prefixed files during migration
When both root/@fix_plan.md and .ralph/@fix_plan.md exist, the root file
now takes priority and the .ralph/@fix_plan.md is removed (backup exists).
This prevents orphaned legacy files after migration.
---------
Co-authored-by: Test User <test@example.com>
When wizard prompt functions (prompt_text, prompt_number, select_option,
select_with_default) were used with command substitution, ANSI-colored
prompts were captured along with user responses, corrupting .ralphrc
values like: PROJECT_NAME="[0;36mProject name[0m [value]: value"
Fix:
- Redirect all display output (colored prompts, validation messages) to
stderr using >&2, matching the pattern already used by select_multiple()
- Keep only the actual response/result on stdout for command substitution
Changes:
- lib/wizard_utils.sh: Add >&2 to prompt_text (lines 84-87),
prompt_number (lines 118-121, 131-151), confirm (lines 44, 60),
select_option (lines 190-215), select_with_default (lines 330-364)
- tests/unit/test_wizard_utils.bats: Add 20 new tests for stdout/stderr
separation and clean command substitution results
- CLAUDE.md: Update test count to 420
Test count: 420 (up from 396)
Co-authored-by: Test User <test@example.com>
* feat(enable): add ralph-enable wizard for existing projects (v0.11.0)
Add interactive wizard and CI version for enabling Ralph in existing projects.
New commands:
- ralph-enable: Interactive 5-phase wizard for humans
- ralph-enable-ci: Non-interactive version with JSON output for CI/automation
New library components:
- lib/enable_core.sh: Shared logic for idempotency, project detection, templates
- lib/wizard_utils.sh: Interactive prompt utilities
- lib/task_sources.sh: Task import from beads, GitHub Issues, PRD documents
Features:
- Auto-detects project type (TypeScript, Python, Rust, Go)
- Auto-detects framework (Next.js, FastAPI, Django, Express)
- Imports tasks from beads, GitHub Issues, or PRD documents
- Generates .ralphrc project configuration file
- Idempotent: safe to run multiple times, respects existing files
- Exit codes: 0 (success), 1 (error), 2 (already enabled)
Updated:
- install.sh: Added new commands to global installation
- ralph_loop.sh: Loads .ralphrc configuration at startup
Tests: 75 new tests (30 enable_core + 23 task_sources + 22 integration)
Total: 396 tests passing (100% pass rate)
Closes#85, #121, #64, #87, #99
* fix(enable): address code review feedback
Fixes from PR #124 review:
1. sed -i portability (ralph_enable.sh:456)
- Use portable sed + mv pattern instead of GNU-only sed -i
2. sed regex portability (lib/task_sources.sh)
- Replace \s with POSIX [[:space:]] character class
- Add sed -E flag for extended regex
3. jq availability check (ralph_enable_ci.sh:177)
- Add check for jq when --json flag is used
4. Unused filter parameter (lib/task_sources.sh:44)
- Pass filter to bd list --filter command
5. Word-splitting in select_multiple (ralph_enable.sh:322)
- Return comma-separated indices instead of space-separated text
- Update caller to parse indices correctly
6. Missing || true for check_existing_ralph (ralph_enable.sh:185)
- Prevent set -e from exiting on non-zero return
7. select_multiple stdout corruption (lib/wizard_utils.sh)
- Redirect interactive output to stderr
- Only final result goes to stdout
8. Color variables not exported (lib/wizard_utils.sh:12)
- Export WIZARD_* color variables for subshells
9. select_option infinite loop (lib/wizard_utils.sh:179)
- Add guard for empty options array
* fix(tests): add missing mocks and exports for new enable feature
- Add RESPONSE_ANALYSIS_FILE export to test_session_continuity.bats setup
- Add mock ralph_enable.sh and ralph_enable_ci.sh to test_installation.bats
- Add mock lib files: enable_core.sh, wizard_utils.sh, task_sources.sh, timeout_utils.sh
All 396 tests now pass.
* fix(config): fix critical issues from PR review
1. .ralphrc Configuration Loading Fix:
- Captured env var state BEFORE setting defaults with _env_* variables
- load_ralphrc now only restores values explicitly set by environment
- .ralphrc settings are now properly applied (not overwritten by defaults)
2. sed Command Injection Fix:
- Replaced sed with awk for .ralphrc updates in ralph_enable.sh
- awk -v pattern safely handles user input without shell injection risk
3. Shell Injection Fix in safe_create_file():
- Replaced echo with printf '%s\n' for safer content handling
- Prevents issues with backslashes, -n, and special characters
4. Specific Error Codes:
- Added ENABLE_INVALID_ARGS=3 for argument errors
- Added ENABLE_FILE_NOT_FOUND=4 for missing files
- Added ENABLE_DEPENDENCY_MISSING=5 for missing deps (e.g., jq)
- Added ENABLE_PERMISSION_DENIED=6 for permission errors
- Updated ralph_enable.sh and ralph_enable_ci.sh to use specific codes
5. Added tests for .ralphrc loading pattern verification
Test count: 398 (up from 396)
* fix(enable): make --force flag actually overwrite existing files
The --force flag was accepted but safe_create_file() always skipped
existing files regardless of ENABLE_FORCE value.
Changes:
- safe_create_file() now checks ENABLE_FORCE environment variable
- When ENABLE_FORCE="true", overwrites existing files instead of skipping
- Added proper logging for overwrite operations
Added tests:
- Verify enable_ralph_in_directory actually changes file contents with --force
- Test safe_create_file overwrites when ENABLE_FORCE is true
- Test safe_create_file skips when ENABLE_FORCE is false
Test count: 400 (up from 398)
---------
Co-authored-by: Test User <test@example.com>
Root cause: Stale completion indicators in .exit_signals and .response_analysis
files persisted across sessions, causing premature exit when combined with
normal completion indicator increments.
Changes:
- Enhanced reset_session() to clear .exit_signals file (resets to empty structure)
- Enhanced reset_session() to remove .response_analysis file
- Session reset now comprehensively clears all exit-related state
Added 2 new tests:
- reset_session clears exit_signals file to prevent premature exit
- reset_session prevents issue #91 scenario (stale completion indicators)
Test count: 321 (up from 319)
Fixes#91
Merged origin/main into PR branch and applied review feedback:
Review fixes:
- Guard against empty result_obj if jq fails (Macroscope)
- Prioritize result object's session_id over init message (CodeRabbit)
- Add regression test for arrays with session_id only in result element
Resolved conflicts:
- CLAUDE.md: Updated test count to 319
All 319 tests pass.
Claude Code CLI outputs a JSON array instead of a single object:
[{type: "system", ...}, {type: "assistant", ...}, {type: "result", ...}]
This caused parse_json_response to fail with "jq: invalid JSON text"
because it assumed the top-level JSON was an object.
Changes:
- Detect if JSON is an array before parsing
- Extract the "result" type message from the array
- Preserve session_id from init message for continuity
- Normalize to object format for existing parsing logic
- Clean up temporary file after processing
Fixes#112
Co-Authored-By: Claude <noreply@anthropic.com>
* feat(structure): migrate Ralph files to .ralph/ subfolder
BREAKING CHANGE: Ralph configuration files now live in .ralph/ subfolder
This refactoring moves all Ralph-specific files into a hidden .ralph/
directory while keeping src/ at the project root. This improves
compatibility with existing tooling and keeps the project root clean.
Changes:
- Move PROMPT.md, @fix_plan.md, @AGENT.md to .ralph/
- Move specs/, logs/, docs/generated/, examples/ to .ralph/
- Move state files (.response_analysis, .circuit_breaker_state, etc.) to .ralph/
- Keep src/ at project root (unchanged)
- Add RALPH_DIR=".ralph" configuration variable
- Add ralph-migrate command for existing projects
- Create migrate_to_ralph_folder.sh migration script
- Update all path references in scripts and tests
- Update documentation (README.md, CLAUDE.md)
New project structure:
project/
├── .ralph/ # Ralph configuration
│ ├── PROMPT.md
│ ├── @fix_plan.md
│ ├── @AGENT.md
│ ├── specs/
│ ├── logs/
│ └── docs/generated/
└── src/ # Source code (unchanged)
Migration: Run `ralph-migrate` in existing projects to upgrade.
All 310 tests pass (100% pass rate).
* chore: add .claude/settings.local.json to .gitignore
* fix: address code review feedback for .ralph/ subfolder structure
Fixes multiple path-related issues identified in code review:
Test fixes:
- Fix create_sample_prompt to use $RALPH_DIR/PROMPT.md in test_session_continuity.bats
- Fix result_file path to use $RALPH_DIR/.json_parse_result in test_json_parsing.bats
- Fix @fix_plan.md and .response_analysis paths in test_cli_modern.bats
- Update templates directory missing test to account for global fallback
Template fix:
- Fix @fix_plan.md reference in templates/PROMPT.md to use .ralph/ prefix
Script fixes:
- Fix PROMPT_FILE comparison in ralph_loop.sh to use $RALPH_DIR/PROMPT.md
- Fix examples migration logic in migrate_to_ralph_folder.sh (remove premature mkdir)
- Move templates directory check AFTER cd in setup.sh (was checking wrong location)
- Add template directory validation with fallback to global templates
All 310 tests pass.
* Update migrate_to_ralph_folder.sh
Co-authored-by: macroscopeapp[bot] <170038800+macroscopeapp[bot]@users.noreply.github.com>
* fix: address code review feedback for .ralph/ subfolder structure
Code Review Fixes:
- Fix test_json_parsing.bats: all result_file and session file paths now use $RALPH_DIR prefix
- Fix ralph_loop.sh help text: paths now show .ralph/.ralph_session, .ralph/.call_count, etc.
- Fix migrate_to_ralph_folder.sh:
- Proper error handling for date command (separate local declaration)
- Use cp -a source/. dest/ pattern to preserve dotfiles and attributes
- Remove 2>/dev/null suppression to surface copy errors
- Update create_files.sh to use .ralph/ structure for embedded scripts
- Update .gitignore with all .ralph/ state file paths
- Add old structure detection in ralph_loop.sh with helpful migration message
Version Update:
- Bump to v0.10.0 (breaking change: structural reorganization)
- Update README.md and CLAUDE.md with new version and release notes
- Add ralph-migrate documentation to Key Commands section
All 310 tests pass.
---------
Co-authored-by: macroscopeapp[bot] <170038800+macroscopeapp[bot]@users.noreply.github.com>
* feat(timeout): add cross-platform timeout support for macOS
Add portable timeout wrapper that automatically detects and uses the
appropriate timeout command based on the platform:
- Linux: Uses standard GNU `timeout` from coreutils
- macOS: Uses `gtimeout` from Homebrew coreutils
Changes:
- Add lib/timeout_utils.sh with detect_timeout_command() and
portable_timeout() functions
- Update ralph_loop.sh to source timeout_utils.sh and use
portable_timeout for Claude Code execution
- Update install.sh to check for coreutils on macOS and provide
installation instructions
- Update test mocks to include gtimeout and portable_timeout
- Update README.md with macOS coreutils installation instructions
- Update CLAUDE.md with timeout_utils.sh documentation
Users on macOS now need to install coreutils: brew install coreutils
* Update model reference in opencode-review workflow
* Update model name in opencode-review workflow
* Update model version in opencode-review workflow
* Update model version in opencode-review workflow
* Update lib/timeout_utils.sh
Co-authored-by: macroscopeapp[bot] <170038800+macroscopeapp[bot]@users.noreply.github.com>
* Update opencode-review.yml
---------
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: macroscopeapp[bot] <170038800+macroscopeapp[bot]@users.noreply.github.com>
Mock claude commands in test_prd_import.bats were hanging on test 240
because they started with 'cat > /dev/null' which blocks waiting for stdin.
When ralph_import.sh calls 'claude --version' during dependency check,
no stdin is provided, causing indefinite hang.
Changes:
- Added --version flag handler to all 7 mock claude command generators
- Handler returns "Claude Code CLI version 2.0.80" and exits before stdin read
- Fixes hang in test 240 and ensures all 33 prd_import tests pass
This fix allows the quickstart installation to complete successfully
without test stalls.
Fixes #N/A (discovered during quickstart installation)
* Reapply "feat(session): implement session expiration with configurable timeout (#83)"
This reverts commit 1ba55a4b9c.
* fix(session): address code review feedback
- Fix integer overflow: return -1 from get_session_file_age_hours on stat
failure instead of 0, preventing false expiration
- Handle stat failure in init_claude_session with WARN log
- Add comprehensive documentation for return values and expiration strategy
- Add 6 behavioral integration tests that verify actual functionality
- Add inline comments explaining 24-hour default rationale
Test count: 286 → 292 (100% pass rate)
* fix(test): use grep-based verification to fix CI failures
Tests that sourced ralph_loop.sh with --help flag failed in GitHub
Actions due to BATS environment differences. Changed behavioral tests
to grep-based code verification that checks implementation patterns
exist without executing the script.
* fix(test): guard main with BASH_SOURCE for safe sourcing
- Add BASH_SOURCE check to only execute main when script is run directly
- Update tests to source script without --help flag
- Convert grep-based verification tests back to functional tests
- Fixes CI failures caused by script execution during sourcing
---------
Co-authored-by: Test User <test@example.com>
- Add CLAUDE_SESSION_EXPIRY_HOURS configuration variable (default: 24)
- Add get_session_file_age_hours() helper with cross-platform stat support
- Modify init_claude_session() to check session age and remove expired sessions
- Add --session-expiry CLI flag to configure expiration (positive integers only)
- Update help text with new option and example
- Add 10 new tests for session expiration (TDD approach)
Closes#51
Test count: 276 → 286 (100% pass rate)
Co-authored-by: Test User <test@example.com>
- Add --output-format json flag for structured Claude CLI responses
- Implement detect_response_format() for JSON vs text detection
- Implement parse_conversion_response() for extracting JSON fields
- Add check_claude_version() for modern CLI feature detection
- Enhance error handling with structured JSON error messages
- Improve file verification with JSON-derived status information
- Maintain backward compatibility with automatic text fallback
- Add 11 new TDD tests for modern CLI features (tests 23-33)
- Update README.md with Modern CLI Features section
- Update CLAUDE.md with v0.9.8 release notes
Test count: 276 (up from 265)
- Fix BSD date parsing to handle milliseconds in ISO timestamps
(e.g., 2026-01-09T10:30:00.123+00:00)
- Document error_count mapping behavior when only has_errors=true is present
- Remove unused has_session_id_field variable
- Add debug logging for session persistence (controlled by VERBOSE_PROGRESS)
- Standardize session filename to .claude_session_id across all files
All 239 tests passing.
- Extend parse_json_response() to support both flat and Claude CLI formats
- Extract result, sessionId, and metadata fields
- Support metadata.files_changed, metadata.has_errors, completion_status
- Parse progress_indicators array for confidence boosting
- Add session management functions for continuity tracking:
- store_session_id(): Persist session with ISO timestamp
- get_last_session_id(): Retrieve stored session ID
- should_resume_session(): Check session validity (24-hour expiration)
- Add get_epoch_seconds() to date_utils.sh for cross-platform epoch time
- Auto-persist sessionId to .session_id file during response analysis
- Add 16 new TDD tests for Claude CLI format and session management
- Update documentation for v0.9.6 (239 tests total)
Test count: 239 (up from 223)
- Remove unused 'load mocks' (mocks.bash not needed)
- Use GIT_AUTHOR_*/GIT_COMMITTER_* env vars instead of git config --global
- Prefix git commands with 'command' to bypass shell function overrides
- Fix tautological assertions in edge case tests:
- Rename test to "succeeds when run in existing directory (idempotent)"
- Assert success ($status -eq 0) instead of always-true condition
README.md:
- Update version badge to v0.9.3
- Update test count to 165 in all locations
- Update test coverage breakdown (111 unit + 54 integration)
test_installation.bats:
- Add missing mock setup.sh in setup() function
- Fix dependency test to mock all three deps (jq, git, node/npx)
- Remove unused source_install_functions helper function
- Add test_installation.bats with full coverage of install.sh
- Tests cover directory creation, command installation, permissions
- Template and lib file copying verification
- Dependency detection with mocked failures (jq, git, node)
- PATH detection and warning system tests
- Uninstallation cleanup verification
- Idempotency testing (run twice without errors)
- End-to-end installation workflow validation
- All tests use isolated temp directories for safety
- Update CLAUDE.md with new test count (165 total)
- Fix npm test script to run tests recursively
- Version bump to v0.9.3
The build_claude_command() function was incorrectly using --prompt-file
which doesn't exist in Claude Code CLI. This fix:
- Replaces --prompt-file with -p flag plus prompt content
- Reads prompt content via $(cat "$prompt_file") before execution
- Adds error handling for missing prompt files
- Maintains shell injection safety through array-based command building
- Updates comments to reflect the correct approach
Adds 6 TDD tests verifying the fix:
- Uses -p flag instead of --prompt-file
- Reads prompt file content correctly
- Handles missing prompt file
- Includes all modern CLI flags
- Handles multiline prompt content
- Prevents shell injection
Test count: 145 -> 151 (all passing)
Address code review finding by adding dedicated test for
--allowed-tools flag validation.
Add code review report documenting:
- 0 critical issues
- 0 major issues
- 1 minor issue (addressed in this commit)
- 6 positive findings
Test count: 27 CLI parsing tests (105 total unit tests)
Refs #10
Add 26 new BATS tests validating all CLI flags in ralph_loop.sh:
- Help flag tests (2): --help, -h short flag
- Flag value tests (6): --calls, --prompt, --monitor, --verbose, --timeout
- Status flag tests (2): --status with/without status.json
- Circuit breaker tests (2): --reset-circuit, --circuit-status
- Invalid input tests (3): unknown flag, invalid timeout, invalid format
- Multiple flags tests (3): combinations, all flags, early exit
- Flag order tests (2): verify order independence
- Short flag tests (6): -c, -p, -s, -m, -v, -t equivalence
Test strategy uses --help as early-exit escape to validate parsing
without triggering main loop execution.
Closes#10
Implements Issue #28 - modernize CLI commands for better Claude integration.
Key changes:
- Add JSON output format support with --output-format flag (default: json)
- Add session continuity with --continue flag and .claude_session_id file
- Add tool permissions via --allowed-tools flag
- Add build_loop_context() for loop-aware context injection
- Add detect_output_format() and parse_json_response() for JSON parsing
- Maintain backward compatibility with text output fallback
- Add version checking with check_claude_version()
New CLI options:
- --output-format json|text: Control Claude output format
- --allowed-tools "Write,Read,Bash(git *)": Restrict tool permissions
- --no-continue: Disable session continuity
Test coverage:
- 20 new JSON parsing tests (test_json_parsing.bats)
- 23 new CLI modern tests (test_cli_modern.bats)
- All 98 tests passing (100% pass rate)
Addresses CodeRabbit outside diff range comment on lines 268-283:
CRITICAL BUG:
The detect_stuck_loop function had a multi-line string handling bug where
only the first error line was checked against historical outputs when
multiple distinct errors were present.
Problem:
grep -q "$current_errors" file.log
# When $current_errors has multiple lines, grep only matches first line
Impact:
Stuck loops with multiple recurring errors would not be detected correctly,
potentially allowing Ralph to continue running despite being genuinely stuck.
Fix:
Changed from simple grep to nested loop checking:
- For each historical output file
- For each error line in current output
- ALL error lines must appear in that file
- Only returns "stuck" if ALL files contain ALL current errors
Used grep -qF for literal fixed-string matching (not regex) to avoid
edge cases with special characters in error messages.
Test Coverage:
Added 2 new test scenarios (7 → 9 total tests):
Test 8: Multiple distinct errors where ALL repeat across history
Current: Error: Build failed + Fatal: DB lost + Exception: NPE
History: All 3 files contain all 3 errors
Expected: Stuck detected ✅
Test 9: Multiple errors where only some repeat
Current: Error: Build failed + Fatal: DB lost
History: Only first error appears consistently
Expected: Not stuck ✅
All existing tests continue to pass, validating backward compatibility.
Test results:
✓ Error detection tests: 13/13 passing
✓ Stuck loop tests: 9/9 passing (was 7/7)
✓ Total: 22/22 tests passing
This ensures detect_stuck_loop correctly handles the real-world scenario
where Ralph gets stuck on multiple simultaneous recurring errors.
Addresses CodeRabbit review comment:
- Outside diff range (lines 268-283): Multi-line error matching bug
Addresses additional CodeRabbit findings in lib/response_analyzer.sh:
CRITICAL (line 267):
- Fixed detect_stuck_loop() to use two-stage filtering
- Was using naive grep -i "error\|failed" pattern
- Now filters JSON fields before extracting errors
- Pattern aligned with analyze_response() for consistency
DEAD CODE (line 18):
- Removed unused STUCK_INDICATORS array
- Array was defined but never referenced in code
- Reduces maintenance burden per coding guidelines
TEST COVERAGE:
- Added comprehensive test suite for detect_stuck_loop()
- New file: tests/test_stuck_loop_detection.sh
- 7 test scenarios validating:
* JSON fields don't trigger false stuck detection
* Actual repeated errors are correctly detected
* Type annotations are properly excluded
* Function returns appropriate exit codes
Test results:
✓ Error detection tests: 13/13 passing
✓ Stuck loop tests: 7/7 passing
✓ Total: 20/20 tests passing
This ensures both error detection functions (analyze_response and
detect_stuck_loop) use identical filtering logic, preventing circuit
breaker false positives across all code paths.
Addresses CodeRabbit review comments:
- Outside diff range comment: line 267 (Critical)
- Outside diff range comment: line 18 (Dead code)
- Outside diff range comment: lines 254-286 (Test coverage)
Aligns error detection patterns across all implementations and improves
test coverage based on CodeRabbit's critical and major findings.
Changes:
1. CRITICAL: Align lib/response_analyzer.sh pattern with ralph_loop.sh
- Changed Stage 1 filter from '"[^"]*\(error\|failed\)"[^"]*":' to '"[^"]*error[^"]*":'
- Removed bare 'cannot' and 'unable to' from Stage 2 (prevent false positives in prose)
- Both files now use identical patterns for consistency
2. MAJOR: Improved test coverage
- Renamed test 10 from "Cannot/unable in error context" to "Error prefix with descriptive message"
- Added test 10a to validate bare "cannot/unable" DON'T trigger false positives
- Now testing 13 scenarios (was 12)
3. MINOR: Added comprehensive test strategy documentation
- Header comments explain two-stage filtering approach
- Documents pattern consistency requirement
- Lists all 13 test scenarios and their purpose
Test results:
✓ All 13 tests passing
✓ Pattern consistency validated across ralph_loop.sh and lib/response_analyzer.sh
✓ False positive scenarios properly excluded
Addresses CodeRabbit review comments:
- r2655862688 (Critical pattern inconsistency)
- r2655862689 (Major test coverage gap)
- r2655862690 (Minor misleading test name)
Fixes circuit breaker opening prematurely due to naive error pattern matching
that treated JSON field names like "is_error": false as actual errors.
Changes:
- ralph_loop.sh: Implement two-stage error detection with JSON filtering
- lib/response_analyzer.sh: Apply same filtering to error counting
- tests/test_error_detection.sh: Add comprehensive test suite (12 scenarios)
Error detection now:
- Filters out JSON field patterns before searching for errors
- Uses context-specific patterns (^Error:, ]: error, Exception, Fatal)
- Avoids type annotations (error: Error) and code identifiers
- Includes debug logging when VERBOSE_PROGRESS=true
Test coverage validates:
✓ JSON fields don't trigger false positives
✓ Real error messages are correctly detected
✓ Mixed content handled properly
✓ Code diffs and documentation excluded
This prevents the consecutive_same_error counter from incrementing on
false positives, eliminating unnecessary circuit breaker trips.