Skip to content

Container Orchestration Engine (Magnum)

Overview

The firstcloud Container Orchestration Engine is built on OpenStack Magnum. Magnum provisions the OpenStack resources required by a cluster.

You can use Magnum through the OpenStack command-line client or UI. The UI does not provide the full cluster lifecycle workflow described here.

Prerequisites

Before creating a cluster, make sure that:

  • The OpenStack CLI is installed and authenticated. See OpenStack CLI.
  • Your project has sufficient compute, network, floating IP, load balancer, and volume quota for the intended cluster size.
  • An SSH key pair exists if the selected template requires one. See SSH key pairs.
  • You have kubectl installed locally to administer the resulting Kubernetes cluster.

Cluster Templates

A cluster template defines the baseline configuration for a cluster. It is important to use a template rather than assuming a particular Kubernetes version or image name, because template availability can differ between regions and projects.

Note

Magnum cluster templates are maintained by firstcolo. The templates available to your project determine the supported Kubernetes versions, images, networking configuration, and optional features.

List the templates available to your project:

openstack coe cluster template list

Inspect a template before using it:

openstack coe cluster template show <template-name>

Pay particular attention to the exposed Kubernetes version, network driver, image, and any labels documented for that template.

Create a Cluster

Warning

To ensure that the cluster does not rely on a single user account, the cluster MUST be created using a service account. This prevents the cluster from being tied to an individual user's account and ensures that access remains available if the user who created the cluster no longer exists.

Create a Kubernetes cluster from an available template. The following example creates a cluster with three worker nodes and associates the my-keypair SSH key pair:

openstack coe cluster create \
  --cluster-template <template-name> \
  --master-count 3 \
  --master-flavor m1.medium \
  --node-count 3 \
  --flavor m1.medium \
  --keypair my-keypair \
  production-kubernetes

Note

All of the above are required fields. Trying to create a cluster without the required fields will result in a cluster in status CREATE_FAILED.

Cluster creation is asynchronous. Check the status until it is CREATE_COMPLETE:

openstack coe cluster show production-kubernetes

Warning

We do not support clusters with an even number of master nodes. By default, a load balancer for the control plane is automatically created, even when the control plane consists of a single node. This ensures a clean cluster upgrade process. The master nodes will always be spread across all availability zones.

Connect to Kubernetes

After the cluster reaches CREATE_COMPLETE, download its kubeconfig:

mkdir -p ~/.kube/firstcloud-production
openstack coe cluster config production-kubernetes --dir ~/.kube/firstcloud-production
export KUBECONFIG=~/.kube/firstcloud-production/config

Verify access to the cluster:

kubectl get nodes
kubectl get pods --all-namespaces

The kubeconfig grants access to the Kubernetes cluster. Store it securely and distribute only the least-privileged credentials required by your users and automation.

Create Worker Node Groups

Create worker node groups using the following command:

openstack coe nodegroup create --node-count 5 production-kubernetes worker-group2
  • You can create the worker node group with a specific flavor using the flag --flavor
  • To create the node group in a specific Availability Zone (e.g. AZ-2), use the flag --labels availability_zone=AZ-2

Scale Worker Nodes

Change the number of worker nodes for a worker node group with the resize command:

openstack coe cluster resize --nodegroup worker-group2 production-kubernetes 7

When scaling down, Kubernetes workloads must be able to tolerate node removal. Drain workloads and verify application replicas, disruption budgets, and persistent-storage requirements before reducing the node count.

Upgrade Clusters

The upgrade paths available to a cluster are determined by its template and driver version. Before planning an upgrade, inspect the cluster and available templates:

openstack coe cluster show production-kubernetes
openstack coe cluster template list

Test upgrades on a non-production cluster first. Confirm application compatibility, Kubernetes version-skew requirements, node image changes, and backup recovery before upgrading a production cluster.

Delete a Cluster

Delete a cluster through Magnum so that Cluster API and Magnum can remove all resources they manage:

openstack coe cluster delete production-kubernetes

Warning

Deleting a cluster removes its Kubernetes control plane and worker nodes. Back up application data and verify the retention policy for persistent volumes before deletion.

Backing up the Kubernetes etcd Database

To create an etcd snapshot from your cluster, connect to a control-plane node using the SSH key configured for the cluster:

ssh <ssh-user>@<control-plane-node>

Control-plane nodes are not exposed to the internet. If you do not have private network access to the cluster, create a jump host with a floating IP and connect to the control-plane node through it.

Run etcdctl with the certificates used by the kubeadm local etcd instance. The snapshot is written to the control-plane node:

ETCD_CONTAINER=$(sudo crictl ps --name etcd --quiet)
SNAPSHOT_DATE=$(date +%F)

sudo crictl exec "$ETCD_CONTAINER" etcdctl \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key \
  snapshot save "/var/lib/etcd/etcd-snapshot-${SNAPSHOT_DATE}.db"

/var/lib/etcd in the container is normally a mount of the control-plane host's etcd data directory, so the snapshot should also be available on the host at:

sudo ls -lh "/var/lib/etcd/etcd-snapshot-${SNAPSHOT_DATE}.db"

Verify that the snapshot is valid:

sudo crictl exec "$ETCD_CONTAINER" etcdutl \
  snapshot status "/var/lib/etcd/etcd-snapshot-${SNAPSHOT_DATE}.db" \
  --write-out=table

Copy the snapshot to secure storage outside the cluster (for instance to Swift).

Warning

Create snapshots regularly and before cluster upgrades or other control-plane changes. An etcd snapshot contains Kubernetes resources, including Secrets; encrypt it at rest and restrict access accordingly.

Troubleshooting

Cluster creation does not complete

Inspect the cluster status and reason:

openstack coe cluster show <cluster-name>

Typical causes are insufficient project quota, an unavailable image or flavor referenced by the template, networking capacity, or an invalid template label. Check the relevant OpenStack resources in the project and contact support with the cluster ID and status_reason if the cause is not clear.

Kubernetes API is unreachable

First confirm that the Magnum cluster is complete and that the kubeconfig was retrieved for the correct cluster. Then check local network access to the API endpoint and verify that any firewall or security-group rules required by the selected template are in place. Do not change driver-managed rules unless instructed by support.

Known Limitations

  • Node groups are limited to 10 worker nodes. To support additional workloads, create additional node groups. This helps ensure proper anti-affinity between the nodes.