I’ve submitted a Flink Improvement Proposal (FLIP) about adding support for the Flink Kubernetes Operator to run Flink’s MiniCluster in a single pod for low-throughput jobs that require isolation.
In the last few weeks, I’ve been working on a proof-of-concept to demonstrate the feasibility of this idea. I’ve done enough to convince myself that this is viable and identify where the issues will be, but I’m looking for community feedback before I take it much further.
Background
The Flink Kubernetes Operator is one of the best ways to run Flink jobs. From the documentation:
Flink deployments are declared like any other Kubernetes workload, and the operator runs their whole operational life:
- Lifecycle Management: deployment, stateful upgrades, rollbacks, and self-healing
- Zero-Downtime Upgrades: blue/green deployments that switch over only once the new version is proven healthy
- Autoscaling: parallelism and memory continuously right-sized to the observed load
- Kubernetes-Native Operations: Helm installation, RBAC, high availability, metrics, logging, and ingress
You create a FlinkDeployment Kubernetes custom resource that points to your Flink application, and the Operator handles provisioning independently schedulable and independently configurable Job Manager Deployment and Task Manager Deployment, configuring them to form a distributed Flink cluster.
This provides scalability and high availability, and is the right approach in a lot of situations. However, this comes at the cost of a fixed baseline cost of at least one JobManager pod and one or more TaskManager pods.
A smaller, lighter-weight alternative would be useful for small or intermittent jobs, where the minimum resource cost of two separately-scheduled pods is disproportionate to the job itself.
A single-pod, self-contained Flink job that starts fast and needs no multi-pod coordination could be a good fit for low-throughput jobs that aren’t suitable for session clusters because they need isolation.