mirror of
https://github.com/langgenius/dify.git
synced 2026-08-28 20:33:21 +08:00
**Problem:**
The telemetry system had unnecessary abstraction layers and bad practices
from the last 3 commits introducing the gateway implementation:
- TelemetryFacade class wrapper around emit() function
- String literals instead of SignalType enum
- Dictionary mapping enum → string instead of enum → enum
- Unnecessary ENTERPRISE_TELEMETRY_GATEWAY_ENABLED feature flag
- Duplicate guard checks scattered across files
- Non-thread-safe TelemetryGateway singleton pattern
- Missing guard in ops_trace_task.py causing RuntimeError spam
**Solution:**
1. Deleted TelemetryFacade - replaced with thin emit() function in core/telemetry/__init__.py
2. Added SignalType enum ('trace' | 'metric_log') to enterprise/telemetry/contracts.py
3. Replaced CASE_TO_TRACE_TASK_NAME dict with CASE_TO_TRACE_TASK: dict[TelemetryCase, TraceTaskName]
4. Deleted is_gateway_enabled() and _emit_legacy() - using existing ENTERPRISE_ENABLED + ENTERPRISE_TELEMETRY_ENABLED instead
5. Extracted _should_drop_ee_only_event() helper to eliminate duplicate checks
6. Moved TelemetryGateway singleton to ext_enterprise_telemetry.py:
- Init once in init_app() for thread-safety
- Access via get_gateway() function
7. Re-added guard to ops_trace_task.py to prevent RuntimeError when EE=OFF but CE tracing enabled
8. Updated 11 caller files to import 'emit as telemetry_emit' instead of 'TelemetryFacade'
**Result:**
- 322 net lines deleted (533 removed, 211 added)
- All 91 tests pass
- Thread-safe singleton pattern
- Cleaner API surface: from TelemetryFacade.emit() to telemetry_emit()
- Proper enum usage throughout
- No RuntimeError spam in EE=OFF + CE=ON scenario
2.5 KiB
2.5 KiB
Task 6: Integration Verification & Diagnostics
Date: 2026-02-05
Diagnostic Implementation
Added operational diagnostics to EnterpriseMetricHandler:
-
Diagnostic Counter Method (
_increment_diagnostic_counter):- Logs diagnostic events at DEBUG level
- Fail-safe: exceptions don't break processing
- Counter names:
enterprise_telemetry.handler.{counter_name} - Labels: optional dict for case-specific tracking
-
Counter Points Added:
deduped_total: Incremented when duplicate events are skippedprocessed_total: Incremented after each case handler (with case label)rehydration_failed_total: Incremented when payload rehydration fails
-
Gateway Logging:
- DEBUG log when gateway is disabled (legacy path)
- DEBUG log for each routing decision (case, signal_type, ce_eligible)
Test Results
- Enterprise telemetry tests: 87/87 PASSED
- Full unit test suite: 4981/4981 PASSED (excluding pre-existing test_event_handlers.py name collision)
- Lint: Clean (ruff)
- Type check: Clean (basedpyright)
Key Patterns
-
Diagnostic Logging Pattern:
def _increment_diagnostic_counter(self, counter_name: str, labels: dict[str, str] | None = None) -> None: try: # Get exporter, log at DEBUG level logger.debug("Diagnostic counter: %s, labels=%s", full_counter_name, labels or {}) except Exception: logger.debug("Failed to increment diagnostic counter: %s", counter_name, exc_info=True) -
Gateway Routing Diagnostics:
logger.debug( "Gateway routing: case=%s, signal_type=%s, ce_eligible=%s", case, route.signal_type, route.ce_eligible, )
Pre-existing Issues Noted
-
Test file name collision:
test_event_handlers.pyexists in both:tests/unit_tests/enterprise/telemetry/tests/unit_tests/core/workflow/graph_engine/event_management/- Workaround: exclude one during test runs
- Not related to this refactor
-
Type annotation issue in
_on_feedback_created:attrs: dictshould beattrs: dict[str, Any]- Pre-existing, not introduced by this task
Verification Checklist
- Diagnostic counters added to metric handler
- DEBUG logging added to gateway
- All telemetry tests pass
- Full unit test suite passes
- Lint clean
- Type check clean
- Feature flag toggle verified (OFF: legacy, ON: gateway)
- No regressions
Next Steps
Ready for production deployment with feature flag control.