Designing Recoverable Background Workflows With Explicit Execution History

  • Thread starter Thread starter Vitaly Raideria
  • Start date Start date
V

Vitaly Raideria

Guest
A failed workflow is rarely a blank slate. Knowing what already happened is the starting point for recovery. Illustration by Raideria; conceptual example, not a product screenshot.

A daily report does not arrive. The process exited with an error. Somewhere before that error, it may have downloaded data, written rows, generated a file, or sent a request to another system.

You can restart it. But what, exactly, are you restarting?

That question is where a background job becomes an operational problem. Starting Python is straightforward. Understanding a partially completed workflow—and deciding which work can safely run again—takes more structure.

An illustrative report pipeline: orders and inventory succeed, report generation fails, and publication waits.


I lead Raideria, the company behind Dagychu, a self-hosted platform for running and operating jobs and pipelines. This article explains the execution model behind it through an illustrative reporting workflow. It is not a customer case study or a performance benchmark.

The useful test for a tool like this is simple: when a job fails, can someone other than its original author understand what happened and choose the next action?

The missing part is the execution history​


Consider a reporting process with four steps:

  1. Fetch orders.
  2. Fetch inventory.
  3. Build a report using both datasets.
  4. Publish the report.

The first two steps can run independently. Report generation needs both. Publication needs a completed report.

Now suppose report generation fails because it encounters an unexpected value. Both downloads have already succeeded. Publication has not started.

An undifferentiated “report failed” status loses information that matters. You need the failed step, its error, the completed work upstream, and the work still waiting downstream. You also need to know whether the inputs from this execution remain available.

Those questions exist whether the original launch mechanism is cron, a queue consumer, a shell command, or an internal service. Changing the trigger alone does not answer them.

For this workflow, I want the recovery boundary to be visible in the same place as the execution history.

Give the workflow a structure you can operate​


Dagychu separates three concepts:

  • A pipeline defines the jobs and their dependencies.
  • A task represents an execution of that pipeline.
  • A job run records an execution attempt of an individual step.

That distinction lets you inspect a failed report step inside a particular task instead of searching for a vaguely related process in a pile of logs.

The dependency portion of our illustrative pipeline could look like this:

Code:
pipeline_name: daily_report
jobs:
  - job_name: fetch_orders
    path: jobs/fetch_orders/main.py
    deps: []

  - job_name: fetch_inventory
    path: jobs/fetch_inventory/main.py
    deps: []

  - job_name: build_report
    path: jobs/build_report/main.py
    deps: [fetch_orders, fetch_inventory]

  - job_name: publish_report
    path: jobs/publish_report/main.py
    deps: [build_report]

This is a dependency sketch, not a complete runnable project. The scripts, project configuration, and input/output mappings are deliberately omitted. A dependency says when a job may run; input mappings say what data it receives.

Dagychu supports declared output keys and downstream inputs that refer to an earlier job's output. For large datasets, I would pass a reference to a durable artifact or a batch identifier, rather than push the entire dataset into execution metadata. That is a workflow design choice, not automatic artifact storage supplied by the orchestrator.

The job boundaries deserve thought. If downloading and publishing are hidden inside one large script, an orchestration layer cannot invent a safe checkpoint between them. Expose the boundaries at which you need to inspect or repeat work.

Keep business logic; make its execution contract explicit​


Adopting an orchestrator does not have to start with rewriting your reporting logic. It does require a clear boundary around how each job accepts inputs, reports outputs, and signals failure.

Dagychu's documented Python job contract uses JSON input and structured JSON output. A nonzero process exit is a failure. The worker records execution state and captures logs.

If your existing logic is already a function, an entry-point script can adapt it to that contract. If it is a command-line tool, an adapter can invoke it and translate its result. Python's subprocess documentation describes the mechanics of invoking commands and checking their exit status.

The important work is not the wrapper itself. It is deciding what success means. A script that catches every exception and exits successfully has hidden the failure from its operator. A script that writes half a dataset before failing needs a recovery policy for those writes.

An explicit execution contract makes these assumptions reviewable.

Dagychu's conceptual topology: browser and scheduler reach the API; the API stores state in PostgreSQL and queues work through RabbitMQ for workers.


The API, state store, queue, and worker have distinct responsibilities. The diagram summarizes core execution components; it is not a complete deployment manifest. Illustration by Raideria.

A rerun is a decision about dependencies and side effects​


In Dagychu, task and job views expose execution state and history, while job detail provides logs and rerun controls. The documented control model includes both a full pipeline rerun and a rerun starting from a selected job.

The second operation needs careful interpretation: it can involve the downstream subgraph. It should not be understood as “only this one script will execute, regardless of dependencies.”

For our report example, the investigation would be:

  1. Open the failed task and identify the failed job.
  2. Read that attempt's logs and error.
  3. Check whether the successful upstream outputs are still valid and accessible.
  4. Fix the cause, then choose the appropriate rerun boundary.
  5. Inspect the resulting execution and confirm the report was actually published.

If only report generation was wrong, and its original inputs remain usable, starting recovery there may be appropriate. If the source snapshot was wrong, recovering from report generation would preserve the wrong inputs. You need to repeat the affected upstream work instead.

This is why a rerun button is useful only alongside execution context.

A recovery decision tree: inspect failure, verify upstream inputs, select the affected upstream boundary or failed step, then check side effects before rerunning.


Select the recovery boundary from the data and side effects, not merely the red status. This is an operator decision guide, not an automatic Dagychu recovery algorithm. Illustration by Raideria.

Retries add another dimension. Dagychu's documented job configuration supports fixed and exponential retry delays. That is useful for failures that may clear without changing the job, such as a temporary dependency outage.

But retrying a publication request after a timeout may publish twice if the first request succeeded remotely. An orchestrator cannot determine the business meaning of that duplicate on its own. The Amazon Builders' Library discussion of idempotent APIs explains this ambiguity and the use of request identifiers to make repeated requests safe.

For a report publisher, an application-level identity might combine the report date and destination. The destination or publishing code must enforce the intended duplicate behavior. Dagychu's control-plane request handling does not turn arbitrary external writes into exactly-once operations.

What self-hosting means here​


Dagychu uses a web interface, an API, a scheduler, workers, PostgreSQL, and RabbitMQ. Workers execute jobs locally or through Docker. Keeping these components in your environment gives you control of their deployment; it also gives you responsibility for their operation.

You still need backups, capacity planning, updates, and appropriate access controls. Docker-based execution also requires a deliberate decision about access to the Docker daemon.

Community is free to self-host and distributed as container images. The public repository contains installation materials and documentation; the application source is not public. Optional product telemetry is enabled by default in Community and can be disabled in Administration. See the repository's licensing and telemetry sections before deployment.

For a first evaluation, focus on a manual demo run and its execution history. Advanced integrations and edition-specific capabilities are separate questions; they are not required for this exercise.

Try one workflow, including its failure path​


The first experiment should be small enough that you understand every side effect. Do not begin by moving your critical overnight processing.

You need Docker Engine and Docker Compose v2. For a fresh installation, clone the distribution repository into a new directory and run its installer:

Code:
git clone https://github.com/raideria-software/dagychu.git
cd dagychu
./install.sh

Open the address printed by the installer and sign in using the admin token it provides. These are first-install commands; use the documented update procedure for an existing deployment.

The distribution seeds demo pipelines. In Administration → Projects, validate and connect the demo project, then open a demo pipeline and create a task. Inspect the job history and logs before changing anything.

For the next test, use a disposable project with a harmless job that deliberately raises an exception. Observe the failed run, remove the intentional failure, and exercise the rerun controls. Verify which jobs run again. This is a suggested evaluation exercise, not a claim that the reporting example above ships as a demo.

Before bringing in real work, you should be able to answer:

  • Can I find the failed step and its error without opening a server terminal?
  • Can I distinguish successful upstream work from work that never started?
  • Do I understand what the rerun action will execute?
  • Can I verify the final business result, beyond a successful process exit?

If your current tooling already answers those questions comfortably, you have a useful baseline. If answering them requires reconstructing the workflow from memory, that is a concrete reason to evaluate Dagychu.

Start with the Dagychu Community installation repository. Run one small workflow, inspect one failure, and recover it. Judge the tool by how clearly you can explain what happened afterward.
 

Thread statistics

Created
Vitaly Raideria,
Replies
0
Views
4
Back
Top