Skip to content

IaaS Network Provider

English | 简体中文

Concepts

  • ENI: Elastic Network Interface
  • Sub-ENI: Secondary Elastic Network Interface
  • VLAN: Virtual Local Area Network

Overview

Spiderpool can integrate with a generic IaaS Network Provider. When Spiderpool allocates or releases Pod IP addresses, it calls the configured provider to bind or unbind the corresponding IaaS-side IP resources on a cloud platform.

Current limitation: IaaS Network Provider mode currently supports Pod IPv4-only allocation. Pod IPv6 and dual-stack provider-mode allocation are not implemented yet.

This feature is useful for public cloud or private cloud environments where an IP address assigned by Spiderpool must also be registered, bound, or programmed in an external cloud network system before the Pod can use it correctly.

Typical use cases include:

  • Allocating auxiliary IP resources from a cloud platform.
  • Binding an IP to a node, ENI, auxiliary network interface, VLAN sub-interface, or other cloud networking resource.
  • Returning cloud-specific attributes such as Pod interface MAC address and VLAN ID to Spiderpool.
  • Releasing the IaaS-side IP binding when Spiderpool releases the Pod IP.

How it works

When the feature is enabled, Spiderpool performs the following calls:

  1. During Pod IP allocation, Spiderpool allocates IPs from Spiderpool IP pools first, then calls the IaaS Network Provider allocation API.
  2. The IaaS Network Provider binds the IP on the cloud platform and returns the cloud-side network attributes.
  3. Spiderpool writes the returned MAC address and VLAN ID into the allocation result, and the VLAN CNI pipeline uses them to configure the Pod interface.
  4. During Pod IP release, Spiderpool calls the IaaS Network Provider release API for each IPv4 address that should be released.
  5. After the IaaS release call returns successfully, Spiderpool releases the IP from the internal IP pool. "Success" here means the IaaS Network Provider has accepted the release request and started the cloud-side cleanup. It does not guarantee that the IaaS-side IP resource is fully released, because the cloud platform may still be processing due to rate limits or asynchronous cleanup.

The IaaS Network Provider is an HTTP service. Spiderpool only defines the API contract and does not depend on a specific cloud vendor implementation.

Usage

Configure the provider URL and HTTP timeout through Helm values:

ipam:
  enableGatewayDetection: false
  enableIPConflictDetection: false
plugins:
  installVlanCNI: true
iaasNetworkProvider:
  serverUrl: "http://iaas-network-provider.iaas-network-provider-system.svc:80"
  httpRequestTimeout: "50s"
spiderpoolController:
  podResourceInject:
    enabled: true
spiderpoolAgent:
  networkResourcePlugin:
    enabled: true
    kubeletRootDir: /var/lib/kubelet
    resourceAdvertisement:
      subENI:
        rules:
          - resourceName: spidernet.io/sub-eni
            defaultMaxCount: 256
            nodeSelector:
              matchLabels:
                key: value
  • If iaasNetworkProvider.serverUrl is empty, Spiderpool does not call the IaaS Network Provider.
  • spiderpoolAgent.networkResourcePlugin.enabled controls Spiderpool network resource advertisement in spiderpool-agent.
  • spiderpoolAgent.networkResourcePlugin.resourceAdvertisement.subENI.rules[].defaultMaxCount is the scheduler-facing total number of auxiliary ENI slots advertised on matching nodes. The example value 256 advertises 256 schedulable resources; Pods that request spidernet.io/sub-eni are constrained by this capacity. Set it to the actual auxiliary ENI capacity available on each node. Helm defaults subENI.rules to an empty list, which disables Sub-ENI advertisement.
  • spiderpoolAgent.networkResourcePlugin.kubeletRootDir controls the kubelet root used to derive the mounted device-plugins and plugins_registry directories. The default is /var/lib/kubelet.
  • spiderpoolController.podResourceInject.enabled controls whether the Pod webhook automatically injects spidernet.io/sub-eni. When set to false, Spiderpool does not add the resource request automatically; users must declare it on Pods to make the scheduler enforce ENI slot capacity.
  • Provider-mode workloads must use IPv4-only Pod IP allocation. Do not enable IaaS Network Provider mode for Pod IPv6 or dual-stack allocation. In those modes, Spiderpool may send IPv6 allocation data to the provider, while the release path currently handles only IPv4 provider resources, which can cause allocation failures or cloud-side resource inconsistency.
  • plugins.installVlanCNI must also be enabled.
  • ipam.enableGatewayDetection and ipam.enableIPConflictDetection must be disabled. This mode is different from the traditional approach of calling CNI first and then calling IPAM. In this mode, IPAM must be called first to obtain the IaaS IP information before calling CNI to complete the Pod network configuration. Therefore, gateway detection and IP conflict detection cannot work in this mode.

Configure the HTTP request timeout

iaasNetworkProvider.httpRequestTimeout controls how long Spiderpool waits for a single provider HTTP call (allocate or release) before treating it as failed.

Provider timing model

A single provider request goes through two stages:

Stage Max duration Description
Rate-limit wait 30 s The provider checks its token bucket. If no slot is available it waits up to 30 s before accepting the request.
Cloud API call 16 s The provider forwards the request to the underlying cloud platform. Network latency and cloud-side processing can take up to 16 s.
Worst-case total ~48 s Sum of the two stages plus a small network round-trip margin.

Setting httpRequestTimeout shorter than ~48 s risks cancelling a request that the provider has already accepted and started executing on the cloud platform. This creates a state inconsistency: Spiderpool treats the call as a failure while the cloud operation may have succeeded or be in progress.

Scenario Recommended httpRequestTimeout
Default / general use 50s (default)
Low-latency private cloud with no rate limiting 20s
High-contention environment with long rate-limit queues 55s–59s (must remain < 100s)

Validation rules

  • Must be a valid Go duration string (e.g. 50s, 1m).
  • Must be greater than 0.
  • Must be less than 2m (static safety limit).
  • Must be less than 100s (the CNI plugin-to-agent timeout for ADD and DEL).
  • Empty or unset defaults to 50s.
  • Validation failure is fatal: the agent and controller will not start with an invalid value.

Time budget hierarchy

Understanding the full budget chain helps explain why httpRequestTimeout has the constraints it does:

Layer Default timeout Description
kubelet sandbox operation 2 min kubelet's default timeout for the entire sandbox setup (Pod network setup). If the CNI pipeline does not complete within this window, the Pod fails to start. This is the outermost budget.
Spiderpool CNI plugin → agent call 100 s The timeout the Spiderpool CNI binary uses when calling the spiderpool-agent over gRPC. This is the budget available to the agent to complete all IPAM and IaaS work before the CNI plugin gives up.
IaaS provider HTTP call 50 s (default) The per-call timeout configured by httpRequestTimeout. Must fit inside the 100 s agent budget alongside all other IPAM work.
Provider worst-case completion ~48 s The maximum time a single provider request can take (30 s rate-limit wait + 16 s cloud API). This is the minimum meaningful value for httpRequestTimeout.

Runtime behavior

Before sending each provider HTTP call, Spiderpool checks how much time remains in the parent CNI operation context (the 100 s agent budget):

  • If the remaining time is less than the provider worst-case (~48 s), Spiderpool does not start the call and returns a parent budget insufficient error immediately. This prevents the provider from consuming a rate-limit slot for a call that cannot complete, which would leave the cloud-side operation in an unknown state.
  • If the remaining time is sufficient, Spiderpool derives a per-call context bounded by httpRequestTimeout. The effective HTTP deadline is min(now + httpRequestTimeout, parent deadline).
  • For each provider call, Spiderpool sends the effective remaining request budget in the X-Request-Timeout-Ms HTTP header. The value is a positive integer in milliseconds, calculated from the request context immediately before the HTTP request is sent. Provider implementations can use this value to bound rate-limit waiting, cloud API calls, and internal retries without relying on clock synchronization with Spiderpool.

Error messages

Message Meaning Suggested action
parent budget insufficient: Xs remaining is less than provider worst-case 48s The CNI pipeline consumed most of the budget before reaching the IaaS call. Check pipeline latency; consider raising the CNI timeout or reducing httpRequestTimeout.
provider-interaction timeout: ... exceeded configured timeout 50s The provider did not respond within httpRequestTimeout. Check provider health; consider raising httpRequestTimeout if provider load is consistently high.
parent budget exhausted: ... cancelled by parent context deadline The parent deadline arrived while the provider was responding. Same as above; the parent budget ran out before the configured timeout.

Note: VLAN-CNI is a VLAN CNI plugin developed by Spiderpool based on the upstream community cni-plugin project. It can be used to integrate with third-party cloud platform IaaS Network Providers, allocating IaaS-layer VLAN network interfaces for containers.

  • VLAN-CNI is a VLAN CNI plugin developed by Spiderpool based on the upstream community cni-plugin project. It can be used to integrate with third-party cloud platform IaaS Network Providers, allocating IaaS-layer VLAN network interfaces for containers.
  • Confirm the maximum number of available ENI slots per node.
  • It is recommended that extension elastic network interfaces on each node do not have IP addresses configured to avoid communication issues caused by inconsistent return paths.

Verify the feature is enabled

After installation, you can verify whether the feature is active by:

  1. Check the ConfigMap
kubectl get configmap spiderpool-conf -n <spiderpool-namespace> -o yaml | grep iaasNetworkProvider

If the output includes iaasNetworkProvider.serverUrl and the value is non-empty, the feature is enabled.

  1. Check agent startup logs
kubectl logs spiderpool-agent-xxx -n <spiderpool-namespace>

Search for IaaS client created successfully in the agent startup logs. If you see this log, the agent has successfully initialized the IaaS client and the feature is active. If you see IaaS provider configuration validation failed, there is a configuration issue; verify that the serverUrl format is correct.

Configure VLAN CNI

When integrating with the IaaS Network Provider, you must use VLAN CNI to create VLAN sub-interfaces for Pods, and configure the VLAN ID and MAC address allocated by the cloud platform on those sub-interfaces. This ensures that the VLAN sub-interface configuration is consistent with the cloud platform, enabling normal network communication.

If the VLAN ID is manually configured at this point, it will be inconsistent with the VLAN ID allocated by the cloud platform, leading to network communication anomalies. Therefore, do not set vlanID in the vlan configuration of SpiderMultusConfig; otherwise vlan-cni will be unable to create a correctly configured VLAN sub-interface for the Pod.

vlan-cni queries the local spiderpool-agent via a Unix socket during Pod creation to obtain the VLAN ID and MAC address allocated from the IaaS, and then creates the VLAN sub-interface in the Pod network namespace based on this information.

Network resource scheduling

Provider-mode workloads can use the Spiderpool device plugin to limit scheduling by auxiliary ENI capacity. The same plugin can also advertise spidernet.io/<master>-nic, allowing workloads to be scheduled only to nodes that have the physical NIC named by the SpiderMultusConfig master field.

Master NIC scheduling is especially useful when interface names differ across node groups. It does not require provider mode. Auxiliary ENI scheduling advertises spidernet.io/sub-eni and is active only when provider mode is enabled.

For master NIC scheduling configuration and troubleshooting, see Spiderpool Device Plugin. The quick start below covers enabling both Sub-ENI count scheduling and master NIC name scheduling in provider mode.

Quick start

The following steps verify spidernet.io/sub-eni capacity scheduling and spidernet.io/<master>-nic name scheduling. Replace the Provider URL, release name, and namespace with values for your environment.

  1. Prepare Helm values

Create iaas-network-provider-values.yaml. It is recommended to configure both Sub-ENI and master NIC resource advertisement so the scheduler constrains placement by auxiliary ENI capacity and by the physical NIC named in the SpiderMultusConfig master field:

iaasNetworkProvider:
  serverUrl: "http://iaas-network-provider.example.svc:80"

spiderpoolController:
  podResourceInject:
    enabled: true

spiderpoolAgent:
  networkResourcePlugin:
    enabled: true
    kubeletRootDir: /var/lib/kubelet
    resourceAdvertisement:
      masterNIC:
        rules:
          - defaultMaxCount: 10000
            nodeSelector:
              kubernetes.io/os: linux
            includeInterfaces:
              - "eth1"
            excludeInterfaces:
              - "eth0"
      subENI:
        rules:
          - resourceName: spidernet.io/sub-eni
            defaultMaxCount: 256
            nodeSelector:
              matchLabels:
                key: value

What the configuration means:

  • iaasNetworkProvider.serverUrl: service address of the IaaS Network Provider.
  • networkResourcePlugin.enabled: enables Spiderpool Device Plugin resource advertisement.
  • masterNIC.rules[]: array of master NIC name resource advertisement rules. Empty rules disable master NIC advertisement.
  • masterNIC.rules[].defaultMaxCount: virtual total capacity advertised for each selected master NIC, default 10000. It only indicates the NIC exists and does not represent bandwidth or Pod limits.
  • masterNIC.rules[].nodeSelector: optional Kubernetes label selector. When set, only matching nodes advertise that master NIC resource. When unset, all nodes are matched. It supports matchLabels and matchExpressions.
  • masterNIC.rules[].includeInterfaces: shell-style glob expressions to select interfaces, e.g. eth*, ens[0-9].
  • masterNIC.rules[].excludeInterfaces: excludes interfaces selected by the same rule; takes precedence over includeInterfaces.
  • subENI.rules[]: array of Sub-ENI resource advertisement rules. Empty rules disable Sub-ENI advertisement.
  • subENI.rules[].defaultMaxCount: default total auxiliary ENI capacity per node.
  • subENI.rules[].nodeSelector: optional Kubernetes label selector. When set, only matching nodes advertise that Sub-ENI resource. It supports matchLabels and matchExpressions.
  • podResourceInject.enabled: allows the webhook to inject spidernet.io/sub-eni and spidernet.io/<master>-nic for eligible Pods automatically.

  • Install or update Spiderpool

helm upgrade spiderpool spiderpool/spiderpool \
  --namespace kube-system \
  --reuse-values \
  --values iaas-network-provider-values.yaml \
  --wait
  1. Verify the installation
kubectl get pod -n kube-system -l app.kubernetes.io/component=spiderpool-agent -o wide
kubectl get nodes -o custom-columns='NAME:.metadata.name,SUB_ENI:.status.allocatable.spidernet\.io/sub-eni,MASTER_NIC:.status.allocatable.spidernet\.io/eth1-nic'

Expected results:

  • With provider mode enabled, matching nodes show SUB_ENI=256 and MASTER_NIC=10000.
  • Nodes that do not satisfy the condition show <none>.

  • Create the SpiderMultusConfig and SpiderIPPool

apiVersion: spiderpool.spidernet.io/v2beta1
kind: SpiderMultusConfig
metadata:
  name: iaas-vlan-config
  namespace: spiderpool
spec:
  cniType: vlan
  vlan:
    master:
      - eth1
    ippools:
      ipv4:
        - pool-eth1
---
apiVersion: spiderpool.spidernet.io/v2beta1
kind: SpiderIPPool
metadata:
  name: pool-eth1
spec:
  gateway: 172.91.0.1
  ips:
    - 172.91.0.100-172.91.0.120
  subnet: 172.91.0.0/24
kubectl apply -f iaas-vlan-config.yaml
  • master must match the interface name selected by masterNIC.rules[].includeInterfaces; in this example, eth1.
  • Do not set vlanID in the vlan configuration; it is allocated dynamically by the IaaS Network Provider.

  • Start a Pod and watch scheduling events

The following example references the VLAN SpiderMultusConfig from the previous step via an annotation. The webhook injects spidernet.io/sub-eni and spidernet.io/eth1-nic automatically:

apiVersion: v1
kind: Pod
metadata:
  name: sub-eni-scheduling
  annotations:
    k8s.v1.cni.cncf.io/networks: spiderpool/iaas-vlan-config
spec:
  containers:
    - name: test
      image: busybox:1.36
      command: ["sh", "-c", "sleep 3600"]
kubectl apply -f sub-eni-pod.yaml
kubectl get events \
  --field-selector involvedObject.kind=Pod,involvedObject.name=sub-eni-scheduling \
  --sort-by=.metadata.creationTimestamp \
  -o custom-columns='TIME:.metadata.creationTimestamp,TYPE:.type,REASON:.reason,MESSAGE:.message' \
  --watch
  1. Verify

When capacity is available, Events show Scheduled. Confirm the Pod status, its node, and the resource requests injected by the webhook:

kubectl get pod sub-eni-scheduling -o wide
kubectl get pod sub-eni-scheduling \
  -o jsonpath='{.spec.containers[0].resources.requests.spidernet\.io/sub-eni}{"\n"}'
kubectl get pod sub-eni-scheduling \
  -o jsonpath='{.spec.containers[0].resources.requests.spidernet\.io/eth1-nic}{"\n"}'

Expected output: sub-eni is 1 and eth1-nic is 1.

Confirm the Pod's node actually advertises the corresponding resources:

NODE_NAME=$(kubectl get pod sub-eni-scheduling -o jsonpath='{.spec.nodeName}')
kubectl get node "${NODE_NAME}" \
  -o jsonpath='{.status.allocatable.spidernet\.io/sub-eni}{"\n"}'
kubectl get node "${NODE_NAME}" \
  -o jsonpath='{.status.allocatable.spidernet\.io/eth1-nic}{"\n"}'

Expected output: sub-eni is 256 and eth1-nic is 10000.

To verify exhaustion behavior, create enough identical Pods for their combined requests to exceed the capacity of all candidate nodes. Excess Pods remain Pending, and Events report FailedScheduling with Insufficient spidernet.io/sub-eni or Insufficient spidernet.io/eth1-nic.

Troubleshooting

  • Confirm iaasNetworkProvider.serverUrl is not empty.
  • Confirm both subENI.rules and masterNIC.rules are not empty.
  • Check defaultMaxCount, nodeSelector, includeInterfaces, and excludeInterfaces.
  • Run ip link show on the target node to confirm the physical NIC named by master exists.
  • If a provider VLAN Pod does not receive sub-eni or <master>-nic, check podResourceInject.enabled, confirm the VLAN SpiderMultusConfig has no vlanID, and verify the Pod references that configuration.

IaaS-side prerequisites

Before running the quick start, platform administrators need to prepare the IaaS side in advance:

  • Create a VPC subnet and bind it to the node's elastic network interface. For example, bind the VPC subnet 172.91.0.0/24 to the physical NIC eth1 on node ECS-01.
  • Confirm the maximum number of auxiliary ENIs that can be bound per node, which is used to set subENI.rules[].defaultMaxCount.

The SpiderMultusConfig and SpiderIPPool created in step 4 of the quick start correspond to the VPC subnet and physical NIC on the IaaS side. Note:

  • master is a required field and must match the physical NIC name on the target node, as well as the NIC selected by masterNIC.rules[].includeInterfaces. Keep the name consistent across candidate nodes, or enable master NIC scheduling to prevent the workload from being placed on nodes that do not provide it.
  • subnet is a required field. It must match the VPC subnet on the cloud platform.

API contract

The provider must implement the following HTTP APIs.

Allocate IPs

Request

POST /v1/apis/network.iaas.io/ipam/allocate-ips
Content-Type: application/json
X-Request-Timeout-Ms: 50000

Request headers:

Header Required Description
X-Request-Timeout-Ms Yes Remaining request budget in milliseconds. The provider should treat this as the maximum time available from when it receives the request, and should return before this budget is exhausted.

Request body:

{
  "podName": "example-pod",
  "podNamespace": "default",
  "podUID": "9f8b7c6d-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "nodeName": "worker-1",
  "iaasIPsAllocationRequest": [
    {
      "ipAddress": "10.0.0.10",
      "subnet": "10.0.0.0/24",
      "parentNicMac": "fa:16:3e:11:22:33"
    }
  ]
}

Fields:

Field Required Description
podName No Pod name.
podNamespace No Pod namespace.
podUID No Pod UID.
nodeName Yes Node where the Pod is scheduled.
iaasIPsAllocationRequest Yes IPs that Spiderpool has allocated and expects the provider to bind.
ipAddress Yes IP address without CIDR prefix.
subnet Yes Subnet CIDR of the IP.
parentNicMac Yes MAC address of the parent NIC that carries the Pod network.

Response

Any HTTP 2xx status code is treated as success.

Response body:

{
  "podName": "example-pod",
  "podNamespace": "default",
  "nodeName": "worker-1",
  "iaasIPsAllocationResponse": [
    {
      "parentNicMac": "fa:16:3e:11:22:33",
      "subnet": "10.0.0.0/24",
      "ipAddress": "10.0.0.10",
      "macAddress": "fa:16:3e:aa:bb:cc",
      "vlanId": 100
    }
  ]
}

Fields:

Field Required Description
iaasIPsAllocationResponse Yes Allocation results returned by the provider.
parentNicMac Yes Parent NIC MAC used by the provider.
subnet Yes Subnet CIDR of the IP.
ipAddress Yes IP address that was bound by the provider.
macAddress No MAC address assigned by the cloud platform for the Pod interface.
vlanId No VLAN ID assigned by the cloud platform.

If macAddress or vlanId is empty, Spiderpool keeps the original allocation result for that field.

Release IP

Request

POST /v1/apis/network.iaas.io/ipam/release-ip
Content-Type: application/json
X-Request-Timeout-Ms: 50000

Request headers:

Header Required Description
X-Request-Timeout-Ms Yes Remaining request budget in milliseconds. The provider should treat this as the maximum time available from when it receives the request, and should return before this budget is exhausted.

Request body:

{
  "podName": "example-pod",
  "podNamespace": "default",
  "podUID": "9f8b7c6d-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "nodeName": "worker-1",
  "parentNicMac": "fa:16:3e:11:22:33",
  "subnet": "10.0.0.0/24",
  "ipAddress": "10.0.0.10"
}

Fields:

Field Required Description
podName No Pod name.
podNamespace No Pod namespace.
podUID No Pod UID.
nodeName Yes Node where the Pod was running.
parentNicMac No Parent NIC MAC. It may be empty in controller-side GC scenarios.
subnet Yes Subnet CIDR of the IP.
ipAddress Yes IP address to release.

Response

The response body is ignored. Any HTTP 2xx status code is treated as success.

Special scenario handling

Allocation must be synchronously successful

Currently, Spiderpool only continues to update the IP status in SpiderIPPool and create or update the SpiderEndpoint object after the Provider has completed the IaaS-side IP binding and returned the network configuration normally.

In some abnormal scenarios:

  • If the Provider or cloud platform throttles the API and the processing takes a long time, causing Spiderpool to time out while waiting for the HTTP response, Spiderpool will treat this allocation as failed.
  • If the Provider side fails to respond, Spiderpool will wait for the timeout period and then treat this allocation as failed.

If the spiderpool-agent does not receive a successful response from the Provider within the configured httpRequestTimeout (default 50s), this allocation will be treated as a failure, and the Pod will be retried according to Kubernetes retry mechanisms.

Release should be idempotent

The release API should be idempotent. If the IP has already been released or does not exist on the cloud platform, the provider should return a 2xx status code when it is safe to consider the IP released.

This avoids repeated CNI DEL or GC retries causing unnecessary failures.

Release may be eventually completed

Some cloud platforms release IaaS IP resources slowly due to cloud-side rate limits or asynchronous cleanup mechanisms. Therefore, IP release may not be fully completed immediately after the provider receives the release request.

Spiderpool requires the provider to accept the release request and start the cloud-side cleanup. The provider should return success when the release request is accepted or when the IP is already released.

Spiderpool calls the IaaS release API before releasing the IP from Spiderpool's internal IP pool. This order avoids re-allocating an IP in Spiderpool before the cloud platform has accepted the release request. If the cloud platform completes the cleanup asynchronously after that, it does not block Spiderpool's IP release flow.

Parent NIC MAC lookup

Spiderpool passes parentNicMac when it can determine the parent NIC MAC address. In agent-side allocation and release, Spiderpool can usually resolve the value from the runtime network environment or cache.

In controller-side GC, Spiderpool may not run in the host network namespace of every node, so it may not be able to resolve the parent NIC MAC. In such cases, Spiderpool may send an empty parentNicMac during release. Provider implementations should tolerate this for the release API.

Abnormal scenario handling

Spiderpool treats the following cases as failures:

  • HTTP request failure.
  • Non-2xx HTTP response status.
  • Invalid allocation response JSON.
  • Allocation response containing unknown IPs.

When release fails, Spiderpool may retry through later cleanup flows depending on where the release is triggered. Provider implementations should therefore make release operations safe to retry.