IaaS Network Provider
English | 简体中文
Concepts
- ENI: Elastic Network Interface
- Sub-ENI: Secondary Elastic Network Interface
- VLAN: Virtual Local Area Network
Overview
Spiderpool can integrate with a generic IaaS Network Provider. When Spiderpool allocates or releases Pod IP addresses, it calls the configured provider to bind or unbind the corresponding IaaS-side IP resources on a cloud platform.
Current limitation: IaaS Network Provider mode currently supports Pod IPv4-only allocation. Pod IPv6 and dual-stack provider-mode allocation are not implemented yet.
This feature is useful for public cloud or private cloud environments where an IP address assigned by Spiderpool must also be registered, bound, or programmed in an external cloud network system before the Pod can use it correctly.
Typical use cases include:
- Allocating auxiliary IP resources from a cloud platform.
- Binding an IP to a node, ENI, auxiliary network interface, VLAN sub-interface, or other cloud networking resource.
- Returning cloud-specific attributes such as Pod interface MAC address and VLAN ID to Spiderpool.
- Releasing the IaaS-side IP binding when Spiderpool releases the Pod IP.
How it works
When the feature is enabled, Spiderpool performs the following calls:
- During Pod IP allocation, Spiderpool allocates IPs from Spiderpool IP pools first, then calls the IaaS Network Provider allocation API.
- The IaaS Network Provider binds the IP on the cloud platform and returns the cloud-side network attributes.
- Spiderpool writes the returned MAC address and VLAN ID into the allocation result, and the VLAN CNI pipeline uses them to configure the Pod interface.
- During Pod IP release, Spiderpool calls the IaaS Network Provider release API for each IPv4 address that should be released.
- After the IaaS release call returns successfully, Spiderpool releases the IP from the internal IP pool. "Success" here means the IaaS Network Provider has accepted the release request and started the cloud-side cleanup. It does not guarantee that the IaaS-side IP resource is fully released, because the cloud platform may still be processing due to rate limits or asynchronous cleanup.
The IaaS Network Provider is an HTTP service. Spiderpool only defines the API contract and does not depend on a specific cloud vendor implementation.
Usage
Configure the provider URL and HTTP timeout through Helm values:
ipam:
enableGatewayDetection: false
enableIPConflictDetection: false
plugins:
installVlanCNI: true
iaasNetworkProvider:
serverUrl: "http://iaas-network-provider.iaas-network-provider-system.svc:80"
httpRequestTimeout: "50s"
spiderpoolController:
podResourceInject:
enabled: true
spiderpoolAgent:
networkResourcePlugin:
enabled: true
kubeletRootDir: /var/lib/kubelet
resourceAdvertisement:
subENI:
rules:
- resourceName: spidernet.io/sub-eni
defaultMaxCount: 256
nodeSelector:
matchLabels:
key: value
- If
iaasNetworkProvider.serverUrlis empty, Spiderpool does not call the IaaS Network Provider. spiderpoolAgent.networkResourcePlugin.enabledcontrols Spiderpool network resource advertisement in spiderpool-agent.spiderpoolAgent.networkResourcePlugin.resourceAdvertisement.subENI.rules[].defaultMaxCountis the scheduler-facing total number of auxiliary ENI slots advertised on matching nodes. The example value256advertises 256 schedulable resources; Pods that requestspidernet.io/sub-eniare constrained by this capacity. Set it to the actual auxiliary ENI capacity available on each node. Helm defaultssubENI.rulesto an empty list, which disables Sub-ENI advertisement.spiderpoolAgent.networkResourcePlugin.kubeletRootDircontrols the kubelet root used to derive the mounteddevice-pluginsandplugins_registrydirectories. The default is/var/lib/kubelet.spiderpoolController.podResourceInject.enabledcontrols whether the Pod webhook automatically injectsspidernet.io/sub-eni. When set tofalse, Spiderpool does not add the resource request automatically; users must declare it on Pods to make the scheduler enforce ENI slot capacity.- Provider-mode workloads must use IPv4-only Pod IP allocation. Do not enable IaaS Network Provider mode for Pod IPv6 or dual-stack allocation. In those modes, Spiderpool may send IPv6 allocation data to the provider, while the release path currently handles only IPv4 provider resources, which can cause allocation failures or cloud-side resource inconsistency.
plugins.installVlanCNImust also be enabled.ipam.enableGatewayDetectionandipam.enableIPConflictDetectionmust be disabled. This mode is different from the traditional approach of calling CNI first and then calling IPAM. In this mode, IPAM must be called first to obtain the IaaS IP information before calling CNI to complete the Pod network configuration. Therefore, gateway detection and IP conflict detection cannot work in this mode.
Configure the HTTP request timeout
iaasNetworkProvider.httpRequestTimeout controls how long Spiderpool waits for a single provider HTTP call (allocate or release) before treating it as failed.
Provider timing model
A single provider request goes through two stages:
| Stage | Max duration | Description |
|---|---|---|
| Rate-limit wait | 30 s | The provider checks its token bucket. If no slot is available it waits up to 30 s before accepting the request. |
| Cloud API call | 16 s | The provider forwards the request to the underlying cloud platform. Network latency and cloud-side processing can take up to 16 s. |
| Worst-case total | ~48 s | Sum of the two stages plus a small network round-trip margin. |
Setting httpRequestTimeout shorter than ~48 s risks cancelling a request that the provider has already accepted and started executing on the cloud platform. This creates a state inconsistency: Spiderpool treats the call as a failure while the cloud operation may have succeeded or be in progress.
Recommended values
| Scenario | Recommended httpRequestTimeout |
|---|---|
| Default / general use | 50s (default) |
| Low-latency private cloud with no rate limiting | 20s |
| High-contention environment with long rate-limit queues | 55s–59s (must remain < 100s) |
Validation rules
- Must be a valid Go duration string (e.g.
50s,1m). - Must be greater than
0. - Must be less than
2m(static safety limit). - Must be less than
100s(the CNI plugin-to-agent timeout for ADD and DEL). - Empty or unset defaults to
50s. - Validation failure is fatal: the agent and controller will not start with an invalid value.
Time budget hierarchy
Understanding the full budget chain helps explain why httpRequestTimeout has the constraints it does:
| Layer | Default timeout | Description |
|---|---|---|
| kubelet sandbox operation | 2 min | kubelet's default timeout for the entire sandbox setup (Pod network setup). If the CNI pipeline does not complete within this window, the Pod fails to start. This is the outermost budget. |
| Spiderpool CNI plugin → agent call | 100 s | The timeout the Spiderpool CNI binary uses when calling the spiderpool-agent over gRPC. This is the budget available to the agent to complete all IPAM and IaaS work before the CNI plugin gives up. |
| IaaS provider HTTP call | 50 s (default) | The per-call timeout configured by httpRequestTimeout. Must fit inside the 100 s agent budget alongside all other IPAM work. |
| Provider worst-case completion | ~48 s | The maximum time a single provider request can take (30 s rate-limit wait + 16 s cloud API). This is the minimum meaningful value for httpRequestTimeout. |
Runtime behavior
Before sending each provider HTTP call, Spiderpool checks how much time remains in the parent CNI operation context (the 100 s agent budget):
- If the remaining time is less than the provider worst-case (~48 s), Spiderpool does not start the call and returns a
parent budget insufficienterror immediately. This prevents the provider from consuming a rate-limit slot for a call that cannot complete, which would leave the cloud-side operation in an unknown state. - If the remaining time is sufficient, Spiderpool derives a per-call context bounded by
httpRequestTimeout. The effective HTTP deadline ismin(now + httpRequestTimeout, parent deadline). - For each provider call, Spiderpool sends the effective remaining request budget in the
X-Request-Timeout-MsHTTP header. The value is a positive integer in milliseconds, calculated from the request context immediately before the HTTP request is sent. Provider implementations can use this value to bound rate-limit waiting, cloud API calls, and internal retries without relying on clock synchronization with Spiderpool.
Error messages
| Message | Meaning | Suggested action |
|---|---|---|
parent budget insufficient: Xs remaining is less than provider worst-case 48s |
The CNI pipeline consumed most of the budget before reaching the IaaS call. | Check pipeline latency; consider raising the CNI timeout or reducing httpRequestTimeout. |
provider-interaction timeout: ... exceeded configured timeout 50s |
The provider did not respond within httpRequestTimeout. |
Check provider health; consider raising httpRequestTimeout if provider load is consistently high. |
parent budget exhausted: ... cancelled by parent context deadline |
The parent deadline arrived while the provider was responding. | Same as above; the parent budget ran out before the configured timeout. |
Note: VLAN-CNI is a VLAN CNI plugin developed by Spiderpool based on the upstream community cni-plugin project. It can be used to integrate with third-party cloud platform IaaS Network Providers, allocating IaaS-layer VLAN network interfaces for containers.
- VLAN-CNI is a VLAN CNI plugin developed by Spiderpool based on the upstream community cni-plugin project. It can be used to integrate with third-party cloud platform IaaS Network Providers, allocating IaaS-layer VLAN network interfaces for containers.
- Confirm the maximum number of available ENI slots per node.
- It is recommended that extension elastic network interfaces on each node do not have IP addresses configured to avoid communication issues caused by inconsistent return paths.
Verify the feature is enabled
After installation, you can verify whether the feature is active by:
- Check the ConfigMap
If the output includes iaasNetworkProvider.serverUrl and the value is non-empty, the feature is enabled.
- Check agent startup logs
Search for IaaS client created successfully in the agent startup logs. If you see this log, the agent has successfully initialized the IaaS client and the feature is active. If you see IaaS provider configuration validation failed, there is a configuration issue; verify that the serverUrl format is correct.
Configure VLAN CNI
When integrating with the IaaS Network Provider, you must use VLAN CNI to create VLAN sub-interfaces for Pods, and configure the VLAN ID and MAC address allocated by the cloud platform on those sub-interfaces. This ensures that the VLAN sub-interface configuration is consistent with the cloud platform, enabling normal network communication.
If the VLAN ID is manually configured at this point, it will be inconsistent with the VLAN ID allocated by the cloud platform, leading to network communication anomalies. Therefore, do not set vlanID in the vlan configuration of SpiderMultusConfig; otherwise vlan-cni will be unable to create a correctly configured VLAN sub-interface for the Pod.
vlan-cni queries the local spiderpool-agent via a Unix socket during Pod creation to obtain the VLAN ID and MAC address allocated from the IaaS, and then creates the VLAN sub-interface in the Pod network namespace based on this information.
Network resource scheduling
Provider-mode workloads can use the Spiderpool device plugin to limit scheduling by auxiliary ENI capacity. The same plugin can also advertise spidernet.io/<master>-nic, allowing workloads to be scheduled only to nodes that have the physical NIC named by the SpiderMultusConfig master field.
Master NIC scheduling is especially useful when interface names differ across node groups. It does not require provider mode. Auxiliary ENI scheduling advertises spidernet.io/sub-eni and is active only when provider mode is enabled.
For master NIC scheduling configuration and troubleshooting, see Spiderpool Device Plugin. The quick start below covers enabling both Sub-ENI count scheduling and master NIC name scheduling in provider mode.
Quick start
The following steps verify spidernet.io/sub-eni capacity scheduling and spidernet.io/<master>-nic name scheduling. Replace the Provider URL, release name, and namespace with values for your environment.
- Prepare Helm values
Create iaas-network-provider-values.yaml. It is recommended to configure both Sub-ENI and master NIC resource advertisement so the scheduler constrains placement by auxiliary ENI capacity and by the physical NIC named in the SpiderMultusConfig master field:
iaasNetworkProvider:
serverUrl: "http://iaas-network-provider.example.svc:80"
spiderpoolController:
podResourceInject:
enabled: true
spiderpoolAgent:
networkResourcePlugin:
enabled: true
kubeletRootDir: /var/lib/kubelet
resourceAdvertisement:
masterNIC:
rules:
- defaultMaxCount: 10000
nodeSelector:
kubernetes.io/os: linux
includeInterfaces:
- "eth1"
excludeInterfaces:
- "eth0"
subENI:
rules:
- resourceName: spidernet.io/sub-eni
defaultMaxCount: 256
nodeSelector:
matchLabels:
key: value
What the configuration means:
iaasNetworkProvider.serverUrl: service address of the IaaS Network Provider.networkResourcePlugin.enabled: enables Spiderpool Device Plugin resource advertisement.masterNIC.rules[]: array of master NIC name resource advertisement rules. Empty rules disable master NIC advertisement.masterNIC.rules[].defaultMaxCount: virtual total capacity advertised for each selected master NIC, default10000. It only indicates the NIC exists and does not represent bandwidth or Pod limits.masterNIC.rules[].nodeSelector: optional Kubernetes label selector. When set, only matching nodes advertise that master NIC resource. When unset, all nodes are matched. It supportsmatchLabelsandmatchExpressions.masterNIC.rules[].includeInterfaces: shell-style glob expressions to select interfaces, e.g.eth*,ens[0-9].masterNIC.rules[].excludeInterfaces: excludes interfaces selected by the same rule; takes precedence overincludeInterfaces.subENI.rules[]: array of Sub-ENI resource advertisement rules. Empty rules disable Sub-ENI advertisement.subENI.rules[].defaultMaxCount: default total auxiliary ENI capacity per node.subENI.rules[].nodeSelector: optional Kubernetes label selector. When set, only matching nodes advertise that Sub-ENI resource. It supportsmatchLabelsandmatchExpressions.-
podResourceInject.enabled: allows the webhook to injectspidernet.io/sub-eniandspidernet.io/<master>-nicfor eligible Pods automatically. -
Install or update Spiderpool
helm upgrade spiderpool spiderpool/spiderpool \
--namespace kube-system \
--reuse-values \
--values iaas-network-provider-values.yaml \
--wait
- Verify the installation
kubectl get pod -n kube-system -l app.kubernetes.io/component=spiderpool-agent -o wide
kubectl get nodes -o custom-columns='NAME:.metadata.name,SUB_ENI:.status.allocatable.spidernet\.io/sub-eni,MASTER_NIC:.status.allocatable.spidernet\.io/eth1-nic'
Expected results:
- With provider mode enabled, matching nodes show
SUB_ENI=256andMASTER_NIC=10000. -
Nodes that do not satisfy the condition show
<none>. -
Create the SpiderMultusConfig and SpiderIPPool
apiVersion: spiderpool.spidernet.io/v2beta1
kind: SpiderMultusConfig
metadata:
name: iaas-vlan-config
namespace: spiderpool
spec:
cniType: vlan
vlan:
master:
- eth1
ippools:
ipv4:
- pool-eth1
---
apiVersion: spiderpool.spidernet.io/v2beta1
kind: SpiderIPPool
metadata:
name: pool-eth1
spec:
gateway: 172.91.0.1
ips:
- 172.91.0.100-172.91.0.120
subnet: 172.91.0.0/24
mastermust match the interface name selected bymasterNIC.rules[].includeInterfaces; in this example,eth1.-
Do not set
vlanIDin thevlanconfiguration; it is allocated dynamically by the IaaS Network Provider. -
Start a Pod and watch scheduling events
The following example references the VLAN SpiderMultusConfig from the previous step via an annotation. The webhook injects spidernet.io/sub-eni and spidernet.io/eth1-nic automatically:
apiVersion: v1
kind: Pod
metadata:
name: sub-eni-scheduling
annotations:
k8s.v1.cni.cncf.io/networks: spiderpool/iaas-vlan-config
spec:
containers:
- name: test
image: busybox:1.36
command: ["sh", "-c", "sleep 3600"]
kubectl apply -f sub-eni-pod.yaml
kubectl get events \
--field-selector involvedObject.kind=Pod,involvedObject.name=sub-eni-scheduling \
--sort-by=.metadata.creationTimestamp \
-o custom-columns='TIME:.metadata.creationTimestamp,TYPE:.type,REASON:.reason,MESSAGE:.message' \
--watch
- Verify
When capacity is available, Events show Scheduled. Confirm the Pod status, its node, and the resource requests injected by the webhook:
kubectl get pod sub-eni-scheduling -o wide
kubectl get pod sub-eni-scheduling \
-o jsonpath='{.spec.containers[0].resources.requests.spidernet\.io/sub-eni}{"\n"}'
kubectl get pod sub-eni-scheduling \
-o jsonpath='{.spec.containers[0].resources.requests.spidernet\.io/eth1-nic}{"\n"}'
Expected output: sub-eni is 1 and eth1-nic is 1.
Confirm the Pod's node actually advertises the corresponding resources:
NODE_NAME=$(kubectl get pod sub-eni-scheduling -o jsonpath='{.spec.nodeName}')
kubectl get node "${NODE_NAME}" \
-o jsonpath='{.status.allocatable.spidernet\.io/sub-eni}{"\n"}'
kubectl get node "${NODE_NAME}" \
-o jsonpath='{.status.allocatable.spidernet\.io/eth1-nic}{"\n"}'
Expected output: sub-eni is 256 and eth1-nic is 10000.
To verify exhaustion behavior, create enough identical Pods for their combined requests to exceed the capacity of all candidate nodes. Excess Pods remain Pending, and Events report FailedScheduling with Insufficient spidernet.io/sub-eni or Insufficient spidernet.io/eth1-nic.
Troubleshooting
- Confirm
iaasNetworkProvider.serverUrlis not empty. - Confirm both
subENI.rulesandmasterNIC.rulesare not empty. - Check
defaultMaxCount,nodeSelector,includeInterfaces, andexcludeInterfaces. - Run
ip link showon the target node to confirm the physical NIC named bymasterexists. - If a provider VLAN Pod does not receive
sub-enior<master>-nic, checkpodResourceInject.enabled, confirm the VLAN SpiderMultusConfig has novlanID, and verify the Pod references that configuration.
IaaS-side prerequisites
Before running the quick start, platform administrators need to prepare the IaaS side in advance:
- Create a VPC subnet and bind it to the node's elastic network interface. For example, bind the VPC subnet
172.91.0.0/24to the physical NICeth1on nodeECS-01. - Confirm the maximum number of auxiliary ENIs that can be bound per node, which is used to set
subENI.rules[].defaultMaxCount.
The SpiderMultusConfig and SpiderIPPool created in step 4 of the quick start correspond to the VPC subnet and physical NIC on the IaaS side. Note:
masteris a required field and must match the physical NIC name on the target node, as well as the NIC selected bymasterNIC.rules[].includeInterfaces. Keep the name consistent across candidate nodes, or enable master NIC scheduling to prevent the workload from being placed on nodes that do not provide it.subnetis a required field. It must match the VPC subnet on the cloud platform.
API contract
The provider must implement the following HTTP APIs.
Allocate IPs
Request
POST /v1/apis/network.iaas.io/ipam/allocate-ips
Content-Type: application/json
X-Request-Timeout-Ms: 50000
Request headers:
| Header | Required | Description |
|---|---|---|
X-Request-Timeout-Ms |
Yes | Remaining request budget in milliseconds. The provider should treat this as the maximum time available from when it receives the request, and should return before this budget is exhausted. |
Request body:
{
"podName": "example-pod",
"podNamespace": "default",
"podUID": "9f8b7c6d-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"nodeName": "worker-1",
"iaasIPsAllocationRequest": [
{
"ipAddress": "10.0.0.10",
"subnet": "10.0.0.0/24",
"parentNicMac": "fa:16:3e:11:22:33"
}
]
}
Fields:
| Field | Required | Description |
|---|---|---|
podName |
No | Pod name. |
podNamespace |
No | Pod namespace. |
podUID |
No | Pod UID. |
nodeName |
Yes | Node where the Pod is scheduled. |
iaasIPsAllocationRequest |
Yes | IPs that Spiderpool has allocated and expects the provider to bind. |
ipAddress |
Yes | IP address without CIDR prefix. |
subnet |
Yes | Subnet CIDR of the IP. |
parentNicMac |
Yes | MAC address of the parent NIC that carries the Pod network. |
Response
Any HTTP 2xx status code is treated as success.
Response body:
{
"podName": "example-pod",
"podNamespace": "default",
"nodeName": "worker-1",
"iaasIPsAllocationResponse": [
{
"parentNicMac": "fa:16:3e:11:22:33",
"subnet": "10.0.0.0/24",
"ipAddress": "10.0.0.10",
"macAddress": "fa:16:3e:aa:bb:cc",
"vlanId": 100
}
]
}
Fields:
| Field | Required | Description |
|---|---|---|
iaasIPsAllocationResponse |
Yes | Allocation results returned by the provider. |
parentNicMac |
Yes | Parent NIC MAC used by the provider. |
subnet |
Yes | Subnet CIDR of the IP. |
ipAddress |
Yes | IP address that was bound by the provider. |
macAddress |
No | MAC address assigned by the cloud platform for the Pod interface. |
vlanId |
No | VLAN ID assigned by the cloud platform. |
If macAddress or vlanId is empty, Spiderpool keeps the original allocation result for that field.
Release IP
Request
POST /v1/apis/network.iaas.io/ipam/release-ip
Content-Type: application/json
X-Request-Timeout-Ms: 50000
Request headers:
| Header | Required | Description |
|---|---|---|
X-Request-Timeout-Ms |
Yes | Remaining request budget in milliseconds. The provider should treat this as the maximum time available from when it receives the request, and should return before this budget is exhausted. |
Request body:
{
"podName": "example-pod",
"podNamespace": "default",
"podUID": "9f8b7c6d-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"nodeName": "worker-1",
"parentNicMac": "fa:16:3e:11:22:33",
"subnet": "10.0.0.0/24",
"ipAddress": "10.0.0.10"
}
Fields:
| Field | Required | Description |
|---|---|---|
podName |
No | Pod name. |
podNamespace |
No | Pod namespace. |
podUID |
No | Pod UID. |
nodeName |
Yes | Node where the Pod was running. |
parentNicMac |
No | Parent NIC MAC. It may be empty in controller-side GC scenarios. |
subnet |
Yes | Subnet CIDR of the IP. |
ipAddress |
Yes | IP address to release. |
Response
The response body is ignored. Any HTTP 2xx status code is treated as success.
Special scenario handling
Allocation must be synchronously successful
Currently, Spiderpool only continues to update the IP status in SpiderIPPool and create or update the SpiderEndpoint object after the Provider has completed the IaaS-side IP binding and returned the network configuration normally.
In some abnormal scenarios:
- If the Provider or cloud platform throttles the API and the processing takes a long time, causing Spiderpool to time out while waiting for the HTTP response, Spiderpool will treat this allocation as failed.
- If the Provider side fails to respond, Spiderpool will wait for the timeout period and then treat this allocation as failed.
If the spiderpool-agent does not receive a successful response from the Provider within the configured httpRequestTimeout (default 50s), this allocation will be treated as a failure, and the Pod will be retried according to Kubernetes retry mechanisms.
Release should be idempotent
The release API should be idempotent. If the IP has already been released or does not exist on the cloud platform, the provider should return a 2xx status code when it is safe to consider the IP released.
This avoids repeated CNI DEL or GC retries causing unnecessary failures.
Release may be eventually completed
Some cloud platforms release IaaS IP resources slowly due to cloud-side rate limits or asynchronous cleanup mechanisms. Therefore, IP release may not be fully completed immediately after the provider receives the release request.
Spiderpool requires the provider to accept the release request and start the cloud-side cleanup. The provider should return success when the release request is accepted or when the IP is already released.
Spiderpool calls the IaaS release API before releasing the IP from Spiderpool's internal IP pool. This order avoids re-allocating an IP in Spiderpool before the cloud platform has accepted the release request. If the cloud platform completes the cleanup asynchronously after that, it does not block Spiderpool's IP release flow.
Parent NIC MAC lookup
Spiderpool passes parentNicMac when it can determine the parent NIC MAC address. In agent-side allocation and release, Spiderpool can usually resolve the value from the runtime network environment or cache.
In controller-side GC, Spiderpool may not run in the host network namespace of every node, so it may not be able to resolve the parent NIC MAC. In such cases, Spiderpool may send an empty parentNicMac during release. Provider implementations should tolerate this for the release API.
Abnormal scenario handling
Spiderpool treats the following cases as failures:
- HTTP request failure.
- Non-
2xxHTTP response status. - Invalid allocation response JSON.
- Allocation response containing unknown IPs.
When release fails, Spiderpool may retry through later cleanup flows depending on where the release is triggered. Provider implementations should therefore make release operations safe to retry.