A Discussion On Infrastructure Sovereignty

A whitepaper published by Palantir reminded me about an issue that is long standing in the world of IT - that is, infrastructure sovereignty. While Palantir scoped their whitepaper on AI, it's an issue that's plagued across infrastructure stacks across the world over a decade.

What is Sovereign Infrastructure?

Up until the early 2000s, every enterprise hosted their own core services - datacenter orchestration, email, communications, database, and more. Active Directory, SCCM, Exchange, SQL Server/Oracle Database/PostgreSQL, Cisco Unified Communications Manager (CUCM), and other on-prem & colocated services drove business. This infrastructure was sovereign - meaning that businesses were responsible for their uptime in exchange for data being resident on their own systems. This made things easily auditable, traceable, and responsibility stayed in the enterprise. If a vendor disappeared tomorrow, business still functioned, and migrations happened on an enterprise's own terms. Data stayed in the enterprise, because everything was under control.

The Defining Timeline

In 2006, Google released Google Apps for Education, a free service for academic institutions for email & calendar. Soon, Google rolled that product to the enterprise, then added GSuite, then more features like identity and Chrome Enterprise to transform it into the Google Workspace we have today. This market also created Office 365 - what was a solution like early GSuite turned into the broader Microsoft 365 we know and despise today. Zoho and others also had their advent here too.

Around the same time, Amazon Web Services had their advent. Soon, Google Cloud Platform and Microsoft Azure joined, then all the other cloud providers, such as Linode, Hetzner, DigitalOcean, and Oracle Cloud Infrastructure had their advent.

Their pitch was this - pay us to serve these for you and then, you don't have to worry about infrastructure. This created the Software as a Service (SaaS) and cloud markets. However, as we all know, SaaS/cloud all really are just someone else's computers & software, creating the fallacy that put technologists in the position we are in today.

The SaaS Fallacy

The quick transition to SaaS put technologists in an uncomfortable position - retain on-prem systems that vendors have effectively put in maintenance mode and not be able to reap security and modern feature benefits while paying the upkeep tax and the increased licensing costs from vendors who use licensing pressure to get you to shift to their SaaS, or shift to SaaS, and reap all the security and feature benefits, at the cost of losing control, paying licensing in perpetuity, and the long term of losing talent to upkeep on-prem, and therefore losing the ability to switch vendors freely.

Most enterprises ended up choosing the latter, which ended up setting the stage for vendors to start spooling down efforts for their on-prem solutions. Exchange Server and Skype for Business Server are effectively in maintenance mode with the move to Subscription Edition. SCCM is still a "just Windows" tool. Jamf Pro on-prem is effectively dead, only available to large accounts, and still requires ancient dependencies (who really wants TomCat in their environment?).

In my professional domain (telephony), it's arguably worse. Cisco is increasingly pressuring enterprises to switch to Webex by making licensing confusing and expensive. Microsoft Teams, Zoom, Five9, and Genesys Cloud have made the unified communications and contact center markets only as a service. SIP is slowly becoming a black box, with Session Border Controllers shifting to SaaS (AudioCodes Live Platform as an example).

In one of my hobby domains (identity and endpoint), Okta, Entra ID, Intune, and Jamf Pro effectively rule the market for identity and MDM. Those are all fully managed SaaS with no escape hatch that's not a full vendor migration.

In cloud compute, developers are increasingly shifting to services locked into the cloud provider. AWS DynamoDB is a fantastic example of that. Kubernetes, while an open standard, set the stage for cloud providers to lock developers into their specific distribution of K8s by providing proprietary load balancers and ingress controllers, and abstracting the control plane nodes.

SaaS Did Not Solve the Availability Problem

The tenancy model ended up showing the industry that availability is arguably harder in the SaaS world. AWS again shows great examples of this - us-east-1 has a big share of catastrophic outages, which plenty of SaaS providers run their compute on, which ends up dragging down most of the internet with it (Genesys Cloud for example).

Non-AWS SaaS isn't exempt. Any Microsoft 365 administrator gets hounded by incidents from M365, ranging from Exchange Online to Teams to Intune. GitHub is the git market leader, yet has outages all the time. Google, while less guilty, still does too.

SaaS Killed Enterprise Control

This is the one I hear about all the time. Microsoft pushes changes on their own schedule - they don't care if they break your environment. Same with every SaaS provider - changes happen on their schedule.

The issue here is compliance and unique enterprise needs. Forced changes require administrators to comply on infrastructure that holds them hostage. Unless you're a customer as large as the Department of War, you don't decide what happens in your environment - the vendor does.

The Stage is Set for a Return to Sovereignty

Cloud compute gave us some of this back by creating and/or strengthening open standards - Kubernetes was born for the cloud. PostgreSQL reached its advent on the cloud. Docker/containerd was born for developers shipping to the cloud. While enterprises ended up losing on the SaaS front, SaaS developers themselves accidentally created the escape mechanism on the infrastructure front.

At the same time, the community saw SaaS sprawl, and are continuing to ship solutions that are the fix to this. Enterprise vendors, especially telephony, have capitalized on the "fall from the cloud" by shipping solutions that can be deployed anywhere. Ribbon Communications is a great example, from SBCs to call control that can be deployed anywhere.

Closed-Source Is Catching Up

Some closed-source solutions have stayed true to serving enterprises with control in mind. In telephony, Ribbon is a great example still. Oracle's Communications division holds in the same way. Mitel has held strong with their on-prem call control, and as much as Cisco pushes to Webex, they still maintain and improve CUCM. Pexip Infinity is my favorite of the bunch - video conferencing, deployable anywhere, and certified for compliance.

Microsoft, while increasingly pushing to the cloud, has realized that they need to keep on-prem around. Exchange Server SE and Skype for Business SE still exist - in maintenance mode. While I fault them for effectively abandoning on-prem, the fact that they keep on-prem around means that they know it's still needed, and that's some credit to give them.

Open-Source Leads the Way Out

The open-source community has done the most work to provide the building blocks for sovereign infrastructure, while staying monetized through enterprise support contracts. A few of my favorite projects:

  • Authentik - an IDP that supports everything from the legacy LDAP & Kerberos stack, to SAML & OIDC, to Shared Signals Framework, and even has device compliance signals in preview (their own agent and Fleet - we'll get to Fleet).
  • Fleet - MDM that focuses on IaC for policy. Supports Apple platforms, Windows, and Linux. Support for Apple DDM, ADE, and User Enrollment. Self-service tools supported. Remediation scripts also supported.
  • Dovecot CE & Dovecot Pro - the premier email stack. Dovecot CE is great to start out with, Dovecot Pro for anything production. Supports IMAP, SMTP, with Oauth2 auth.
  • Proxmox VE - KVM-based hypervisor stack. Clustering supported. Supports full-fat VMs and LXC containers.
  • Talos Linux - an immutable and opinionated OS for K8s, using only declarative configuration. Management tools are fault-tolerant, so that a bad update doesn't kill your cluster.
  • Netbird - Zero-Trust Network Access solution. Cloud and self-hosted options. OIDC supported for auth. BYO relays for self-hosted.

These projects prove that you can do enterprise-grade with open-source, and still get the support that you need. The tools are here.

Cloud Has its Advantages

I will admit, cloud and some SaaS has its place - dismissing some services when others have shifted the market would be unfair when they truly do have their place. The ones I use at home, and would use for commercial production would be:

  • Cloudflare - The internet aaS that we all know and love. I'm a heavy user of their DNS, WAF, Workers, and Zero Trust solutions. I would use it in commercial production because of network footprint (especially when it comes to Workers and ZTNA - both workloads benefit from the edge footprint), developer friendliness, and transparent billing.
  • Oracle Cloud Infrastructure - My main hyperscaler. A true cloud platform - unlike Hetzner, which is really just a bunch of VPSes. VPCs, container services, bastions, managed K8s, Terraform state manager, and more basic cloud products included. Generous free tier, with fair rates (especially when using Ampere A1 instances).

Cloud isn't bad - lock in is. Deploying your workloads in the cloud should be encouraged - with the understanding that its footprint should shrink over time as you move to on-prem & colocated.

How We Move Forward as an Industry

As technologists, we need to bring our knowledge to leadership in our enterprises - there are true financial and operational gains from "falling from the cloud," but it's very unknown as the norm has been set for some time. As most sovereign deployments can be tested without license - do it. As technologists, we know what we do best - go get it done.

← Back to blog