3 minute read

This case study documents a RESTful FastAPI controller for a cleaning robot, with a greedy A*-Manhattan pathfinding algorithm, SQLite history, dockerization, demo, and automated tests. See the GitHub Repository for the full code and documentation.

Situation

In many Machine Learning (ML) systems, once we have developed and validated our models, we expose them via a REST API so other systems can interact with them. We often package this API as a small, independently deployable service for maintainability and scaling. In this project, I built a service for controlling different cleaning robot models from a fictional company, ACME Inc., from environment definition through planning and execution to history tracking and extraction.

Task

My goal was to engineer a production-style end-to-end service, which:

  • Uses A REST API to control the robot’s navigation and actions.
  • Reads different grid-like map formats, such as TXT and JSON.
  • Uses either pre-defined movements or an autonomous planning algorithm to navigate the environment.
  • Tracks information regarding cleaning history and stores it in a persistent database.
  • Supports different robot models with distinct capabilities:
    • A Base model;
    • A Premium model that avoids re-cleaning pre-cleaned tiles to save power.
  • Handles collisions by stopping the cleaning session on impacts with obstacles.
  • Employs a modular design pattern, allowing for separation between the API layer, planning algorithms, robot logic, and history.
  • Provides a reproducible and extensible development and deployment environment using Docker.
  • Includes comprehensive documentation, automated tests with full coverage, and a demo client to showcase functionality.
Example maps for the robot API.
Example maps for the robot API: (left) TXT format, (right) JSON format.

Actions

I structured the solution into five main components:

1. Environment and Robot Models

I constructed an Environment class that ingests both TXT and JSON map formats, validating the structure and enforcing boundaries/obstacles. I then created two robot models, BaseRobot and PremiumRobot. PremiumRobot extends BaseRobot, overriding the cleaning logic to skip already-clean tiles. Both robots are equipped with a collision-detection logic that terminates cleaning and enters an error state when hitting an obstacle.

2. API

I implemented the REST API using FastAPI, allowing for schema-driven development and ensuring that key inputs are validated at the API boundary. The API exposes endpoints for:

  • setting the map: /set-map
  • planning routes between points: /plan
  • planning routes for full coverage: /plan-coverage
  • executing cleaning commands: /clean
  • retrieving cleaning history: /history

3. Planning Algorithms

The service supports three modes of operation:

  • Direct commands, telling the robot the sequence of moves to execute.
  • Point-to-point planning, where the robot calculates a path between two coordinates.
  • Full coverage planning, where the robot autonomously determines a route to clean all reachable tiles.

For point-to-point planning, I implemented an A* (A-Star) pathfinding algorithm with a Manhattan heuristic for 4-connected grid movement.

I then extended this with a greedy coverage strategy to achieve full-room cleaning. The greedy strategy repeatedly selects the nearest unvisited reachable tile. It uses A* to bridge the gap between its current position and the next dirty tile, enabling the robot to calculate a path to clean all reachable tiles autonomously.

Planner route visualisation.
Example of a planned route using the greedy A*-Manhattan algorithm for the Demo Map.

4. Cleaning History

For tracking cleaning session history, I used SQLite with SQLAlchemy, logging each session’s unique identifier, start time, robot model, final state (completed/error), number of actions, number of cleaned tiles, and duration. The history endpoint exports this data as CSV for easy analysis.

5. Reproducibility & Tooling

To ensure the solution is production-style, I implemented several features:

  • I containerised the entire application using Docker/Compose (Gunicorn + Uvicorn) to provide a reproducible runtime.
  • I used uv for deterministic dependency management and just to standardise command-line operations
  • I built a CLI end-to-end demo script (demo.py), allowing full run and route planning simulation without manual API interaction.

Result

The final deliverable is a self-contained environment that passes a comprehensive pytest suite and includes an end-to-end demo script. The application successfully distinguishes between robot models and provides a stable API for control and history.

Scars, Lessons, and Trade-offs

During the implementation, I made several choices:

Choise Reason
State Management (Singleton Trade-off) I hold the map in memory as an app-attached singleton for simplicity and fast (O(1)) collision detection. This made the API stateful, hindering horizontal scaling as it can cause desynchronization in a multi-worker production Kubernetes environment. The solution is moving the environment state to either a shared Redis instance or a session-aware database.
Persistence Layer SQLite via SQLAlchemy is effective for small or prototype projects, but limits write concurrency. For a multi-user production system, I would transition to PostgreSQL.
FastAPI over Flask Asynchronous support, typed schemas, and request validation with Pydantic.
Docker Reproducibility, removing “works on my machine” situations.
Greedy A* While simple and effective, this strategy is sub-optimal and inflexible to environmental changes. A deep reinforcement learning algorithm would provide better adaptability, generalisation and optimal route planning