Site reliability engineering
We are a group of engineers who have been working at the software infrastructure layer for decades. The title varies from place to place: site reliability engineering, devops, or platforms engineering. The goal is the same: to ensure that application software can be deployed and run securely and reliably. We have extensive experience with clouds as well as bare metal, and we specialize in workloads that straddle the boundary between the two. You can hire us to bring this experience to your team.
We can act as your external operations team. We love helping smaller companies, at that point where your needs start to outgrow your initial team, but it’s still too early to start a dedicated in-house ops team. Especially when you have tight uptime requirements, but 24/7 oncall is out of reach for your team alone. Where we really shine though, is when we act as an extension of your existing ops team, and we help you bring your practices to the next level.
We can help you with general platforms engineering topics such as:
- Monitoring and alerting, for example using Prometheus and AlertManager.
- Infrastructure-as-code adoption, for example using OpenTofu or Terraform.
- Configuration management on fleets of Linux and BSD servers, for example using Ansible, PyInfra, or Nix.
- Network configuration, for example setting up a private WireGuard network across cloud and on-premise.
- Configuring and debugging DNS, both on the nameserver and resolver side.
- Adopting service discovery, for example using Consul or CoreDNS.
- Setting up and hardening secrets management, for example using OpenBao or Vault.
- Setting up, configuring, and integrating identity and access management.
- Containerization and Linux namespaces, for example using Podman, Docker, or Systemd.
- Service management, on a single host as well as cluster-wide, for example using Systemd, Docker Compose, Kubernetes, or Nomad.
- Configuring and operating databases such as PostgreSQL and MariaDB.
- Configuring and optimizing build systems, remote builds, continuous integration pipelines, and evolving those towards continuous deployment.
- Achieving reproducible builds.
- Operating open-weight LLMs for privacy-preserving AI assistants and agents.
We specialize in topics around bare metal and digital autonomy:
- Leveraging the upsides of bare metal, such as increased performance and affordable network bandwidth.
- Navigating the challenges of bare metal, being resilient in case of hardware failure, and high availability.
- Hosting with European vendors.
- Self-hosting open source software, both user-facing (e.g. Mattermost or NextCloud) and developer-facing (e.g. Prometheus, Postgres, and Kubernetes). See also our open source adoption page.