Skip to content

Failure Observability

Goal

Distinguish a result, failure trend, and recent diagnostic event. A task exception must not disappear silently even when a caller does not immediately consume its future.

Routing and failure are separate observation paths

Automatic routing explains why a backend was selected, but does not change the failure model. RoutingDecision does not mean work completed or was accepted; failure events do not explain an allowed CPU fallback.

QuestionDefault entryScope
Did this call complete, and what is its result?future.get()One result-bearing task; exception rethrows here
Why was this route selected or rejected?get_last_routing_decision() / routing callbackIntent, fallback, and preflight explanation for submit_auto() / dispatch_auto()
Did a bounded queue accept it?DispatchResult::acceptedOne lock-free or real-time admission; not completion
Did a long-lived worker start or stop?WorkerHandle and worker statusStartup, lifecycle, exit reason; not protocol readiness
What failure just happened in the service?set_failure_callback()Immediately bridge to logs, alerts, or telemetry
How many failures of this type accumulated?get_failure_status()Health checks, dashboards, threshold alerts
What is the context of recent failures?get_recent_failures()Diagnosis, support bundle, bounded history

Failure observation paths

QuestionDefault entryScope
Did this call succeed, and what is its result?future.get()One result-bearing task; exception rethrows here
What failure just happened in the service?set_failure_callback()Immediately bridge to logs, alerts, or telemetry
How many failures of this type accumulated?get_failure_status()Health checks, dashboards, threshold alerts
What is the context of recent failures?get_recent_failures()Diagnosis, support bundle, bounded history

Set a callback after initialization, then retain get() where a result is needed:

cpp
#include <atomic>
#include <exception>
#include <iostream>
#include <stdexcept>

#include <executor/executor.hpp>

int main() {
    auto& executor = executor::Executor::instance();
    std::atomic<int> callbacks{0};
    executor.set_failure_callback([&](const executor::ExecutorFailureEvent&) { ++callbacks; });

    auto failed = executor.submit([]() -> int {
        throw std::runtime_error("expected observability failure");
    });

    try {
        static_cast<void>(failed.get());
    } catch (const std::exception&) {
    }

    executor.wait_for_completion();
    const auto status = executor.get_failure_status();
    const auto recent = executor.get_recent_failures();
    std::cout << "failures=" << status.task_exception_count
              << ", callback=" << callbacks.load()
              << ", recent=" << recent.size() << '\n';

    executor.clear_recent_failures();
    executor.shutdown();
    return status.task_exception_count == 1 && callbacks == 1 && recent.size() == 1 ? 0 : 1;
}
bash
./build/examples/tutorial/tutorial_06_observability
text
failures=1, callback=1, recent=1

future.get() remains the result and exception boundary for one task. Routing decisions, callbacks, counts, and recent events provide distinct explanation or service-level observation; none replaces another. Routing callbacks isolate callback exceptions just like failure callbacks. An allowed CPU fallback keeps fell_back = true and its FallbackPolicy explanation without increasing user-task failure counts.

Recent-event retention

  • get_recent_failures(0) returns the entire current buffer; a positive argument returns that many newest events.
  • set_recent_failure_capacity(n) configures ring-buffer capacity. At 0, events are not retained, but cumulative counts and callback still work.
  • clear_recent_failures() clears only diagnostic history; it does not reset cumulative get_failure_status() counts.

Choose capacity from memory budget and incident investigation window. Do not keep unbounded process-local history.

Callback boundary

The failure callback runs on Executor's failure-recording path. Keep it short and nonblocking, and own any external I/O policy. An exception thrown by the callback is isolated and does not terminate a worker/background thread. For complex handling, enqueue a small event into application logging or alert infrastructure.

Failures are not interchangeable

TaskException, SubmitRejected, WaitTimeout, real-time drops, GPU failure, and safe tuning fallback can all enter ExecutorFailureStatus, but have different meanings. Task exception needs a business-result decision; wait timeout means unfinished work; tuning fallback may still run safely. A routing capability snapshot is not a reservation: stop, a full queue, and object-pool exhaustion still surface through DispatchResult, future rejection, and appropriate failure events. Communication events remain in local executor::comm callbacks/statistics by default.

Next: monitoring and sampling for throughput, success/failure, and execution-time trends; bounded waiting and status for wait timeout decisions.