GitHub Issue: #252 Status: Planning Phase Priority: High Estimated Effort: 3-5 sprints Created: 2025-10-27
AutoBot's backend contains 1,070 exception handlers across 141 files with 2,123 error statements, representing massive code duplication. Only 1 centralized utility (autobot-backend/utils/error_boundaries.py) exists and is barely used (2 imports). This plan outlines a phased approach to centralize error handling, maximize code reuse, and improve maintainability.
- Total Backend Files: 141
- Exception Handlers: 1,070
- Error Statements: 2,123
- Centralized Utilities: 1 (
error_boundaries.py) - Utility Usage: 2 imports (1.4% adoption)
- Code Duplication: ~95%
autobot-backend/utils/error_boundaries.py:
- ✅ Centralized error handling utility
- ✅ Structured error response format
- ✅ HTTP exception wrapping
- ❌ Minimal adoption (only 2 files use it)
- ❌ Limited error categorization
autobot-backend/api/error_monitoring.py:
- ✅ Demonstrates intended pattern
- ❌ Not widely adopted across API endpoints
-
Try-Catch with Logging (65% of handlers)
try: # operation except Exception as e: logger.error(f"Failed to do X: {e}") return {"error": str(e)}
-
HTTPException Raising (25% of handlers)
try: # operation except Exception as e: raise HTTPException(status_code=500, detail=str(e))
-
Silent Exception Swallowing (10% of handlers)
try: # operation except Exception: pass # Or return None
- 95% of error handling code is duplicated
- Same try-catch patterns repeated in 1,070 locations
- Inconsistent error message formats
- Maintenance nightmare (fix once → update 1,070 locations)
- Different endpoints return different error formats
- Some return
{"error": "..."}, others{"detail": "..."} - Frontend must handle multiple error response formats
- Difficult to implement unified error display
- No distinction between transient vs permanent failures
- No retry guidance for clients
- No error severity levels (warning vs critical)
- No error code standardization
- Inconsistent logging levels
- Missing error context (user_id, session_id, operation)
- No error aggregation or trending
- Difficult to debug production issues
- Status codes hardcoded throughout codebase
- Error messages hardcoded in exception handlers
- No centralized error message catalog
Goal: Expand error_boundaries.py into comprehensive error handling framework
Components to Add:
-
Error Categories Enum
class ErrorCategory(Enum): VALIDATION = "validation" # Client error (400) AUTHENTICATION = "authentication" # Auth error (401) AUTHORIZATION = "authorization" # Permission error (403) NOT_FOUND = "not_found" # Resource missing (404) CONFLICT = "conflict" # Resource conflict (409) RATE_LIMIT = "rate_limit" # Too many requests (429) SERVER_ERROR = "server_error" # Internal error (500) SERVICE_UNAVAILABLE = "service_unavailable" # Transient (503) EXTERNAL_SERVICE = "external_service" # Upstream error (502)
-
Standardized Error Response Class
@dataclass class ErrorResponse: category: ErrorCategory message: str code: str # e.g., "KB_001", "AUTH_002" status_code: int details: Optional[Dict[str, Any]] = None retry_after: Optional[int] = None trace_id: Optional[str] = None
-
Context-Aware Error Handler
def handle_error( error: Exception, context: Dict[str, Any], category: ErrorCategory, operation: str ) -> ErrorResponse: """ Centralized error handling with: - Automatic logging with context - Error categorization - Retry guidance - Trace ID generation - Metrics collection """
-
Decorator-Based Error Handling
@with_error_handling( category=ErrorCategory.SERVER_ERROR, operation="knowledge_base_query" ) async def get_facts(session_id: str): # Implementation # Errors automatically caught, logged, and formatted
Goal: Migrate API endpoints to use centralized error handling
Priority Order:
-
High-Traffic Endpoints (sprint 2):
autobot-backend/api/chat.py- Chat API (highest traffic)autobot-backend/api/knowledge.py- Knowledge base APIautobot-backend/api/agents.py- Agent orchestrationautobot-backend/api/web_research_api.py- Research API
-
Medium-Traffic Endpoints (sprint 3):
autobot-backend/api/session_api.py- Session managementautobot-backend/api/workflow_api.py- Workflow executionautobot-backend/api/file_browser.py- File operationsautobot-backend/api/terminal_api.py- Terminal commands
-
Low-Traffic Endpoints (sprint 3):
- Configuration APIs
- Health check endpoints
- Monitoring APIs
Migration Pattern:
# BEFORE (duplicated pattern):
@app.get("/api/facts/{fact_id}")
async def get_fact(fact_id: str):
try:
result = kb.get_fact(fact_id)
return result
except Exception as e:
logger.error(f"Failed to get fact: {e}")
raise HTTPException(status_code=500, detail=str(e))
# AFTER (centralized):
@app.get("/api/facts/{fact_id}")
@with_error_handling(
category=ErrorCategory.SERVER_ERROR,
operation="get_fact"
)
async def get_fact(fact_id: str):
return kb.get_fact(fact_id)
# Errors automatically handled by decoratorGoal: Migrate core services to use error boundaries
Files to Migrate:
src/knowledge_base.py- Knowledge base servicesrc/chat_workflow_manager.py- Workflow managementsrc/llm_interface.py- LLM integrationsrc/agent_orchestrator.py- Agent coordinationsrc/autobot_memory_graph.py- Memory graph operations
Pattern:
- Replace 95% of try-catch blocks with
@with_error_handling - Keep critical 5% for specific error recovery logic
- Add context to all error handlers (session_id, user_id, operation)
Goal: Eliminate hardcoded error messages
Create: config/error_messages.yaml
errors:
KB_001:
category: server_error
message: "Failed to retrieve knowledge base fact"
status_code: 500
retry: true
KB_002:
category: not_found
message: "Knowledge base fact not found"
status_code: 404
retry: false
AUTH_001:
category: authentication
message: "Invalid or expired session"
status_code: 401
retry: falseLoad at Startup:
from src.utils.error_catalog import load_error_catalog
ERROR_CATALOG = load_error_catalog("config/error_messages.yaml")Goal: Add error tracking and alerting
Components:
-
Error Metrics Collection
- Error rate by endpoint
- Error rate by category
- Error response time distribution
- Retry success rate
-
Error Aggregation
- Group similar errors
- Detect error spikes
- Track error trends over time
-
Alerting Rules
- Alert on error rate > threshold
- Alert on new error types
- Alert on cascade failures
-
Error Dashboard
- Real-time error monitoring
- Error distribution charts
- Top failing endpoints
- Target: Reduce duplicated error handling code by 95%
- Target: Centralized error handling adoption > 90%
- Target: Error message hardcodes eliminated (100% in catalog)
- Target: Mean Time To Diagnosis (MTTD) < 5 minutes
- Target: Error response time < 100ms
- Target: Error categorization accuracy > 95%
- Target: New endpoint error handling < 5 lines of code
- Target: Error handling pattern documentation complete
- Target: Error handling test coverage > 80%
- Keep existing error handling patterns during migration
- Add deprecation warnings to old patterns
- Run both systems in parallel during transition
- Gradual cutover endpoint by endpoint
- Unit Tests: Test error boundary functions
- Integration Tests: Test API error responses
- E2E Tests: Test frontend error handling
- Load Tests: Verify error handling performance
- Deploy enhanced error_boundaries.py (no breaking changes)
- Migrate high-traffic endpoints (validate in production)
- Monitor metrics (error rate, response time)
- Migrate remaining endpoints (batch by service area)
- Remove deprecated patterns (final cleanup)
- Mitigation: Maintain backward compatibility during migration
- Mitigation: Feature flags for new error handling
- Mitigation: Gradual rollout with monitoring
- Mitigation: Benchmark error handler performance
- Mitigation: Async error logging to avoid blocking
- Mitigation: Caching for error message lookups
- Mitigation: Automated detection of old patterns
- Mitigation: Pre-commit hooks to enforce new patterns
- Mitigation: Dashboard tracking migration progress
- Implement ErrorCategory enum
- Create ErrorResponse dataclass
- Implement handle_error() function
- Create @with_error_handling decorator
- Add context injection (trace_id, session_id)
- Write unit tests for error boundaries
- Update documentation
- Migrate chat.py endpoints
- Migrate knowledge.py endpoints
- Migrate agents.py endpoints
- Migrate web_research_api.py endpoints
- Migrate session_api.py endpoints
- Migrate workflow_api.py endpoints
- Migrate file_browser.py endpoints
- Migrate terminal_api.py endpoints
- Write integration tests
- Migrate knowledge_base.py
- Migrate chat_workflow_manager.py
- Migrate llm_interface.py
- Migrate agent_orchestrator.py
- Migrate autobot_memory_graph.py
- Write service-level tests
- Create config/error_messages.yaml
- Implement error catalog loader
- Migrate hardcoded messages to catalog
- Add catalog validation in tests
- Document error code conventions
- Implement error metrics collection
- Create error aggregation system
- Setup alerting rules
- Build error dashboard
- Document monitoring setup
autobot-backend/utils/error_boundaries.py- Base utility (exists)config/error_messages.yaml- Error catalog (new)- Pre-commit hooks - Pattern enforcement (new)
- Monitoring dashboard - Observability (new)
- None (pure refactoring, no new libraries)
| Phase | Sprint | Duration | Deliverable |
|---|---|---|---|
| Phase 1 | Sprint 1 | 2 weeks | Enhanced error_boundaries.py |
| Phase 2a | Sprint 2 | 2 weeks | High-traffic API migration |
| Phase 2b | Sprint 3 | 2 weeks | Remaining API migration |
| Phase 3 | Sprint 4 | 2 weeks | Core service migration |
| Phase 4 | Sprint 4 | 1 week | Error message catalog |
| Phase 5 | Sprint 5 | 2 weeks | Monitoring & observability |
Total Duration: 11 weeks (5 sprints)
- Immediate: Review and approve this plan with team
- Sprint 1 Start: Begin Phase 1 implementation
- Before Migration: Establish baseline error metrics
- During Migration: Weekly progress tracking meetings
- Post-Migration: Retrospective and lessons learned
- Current Implementation:
autobot-backend/utils/error_boundaries.py - Example Usage:
autobot-backend/api/error_monitoring.py - Error Statistics: Analysis conducted 2025-10-27
- Zero Hardcode Policy:
CLAUDE.mdSection 🚫
Document Owner: AutoBot Development Team Last Updated: 2025-10-27 Status: Awaiting Approval