Node Discovery
Introduction
Node discovery enables BanyanDB nodes to locate and communicate with each other in a distributed cluster. The discovery registry is consulted in three scenarios:
- Schema Registry Bootstrap: Every node (data and liaison) runs a property schema client that must reach an embedded schema server on a data node to synchronize metadata. The discovery registry is the only source of schema server endpoints, so a misconfigured discovery mode prevents the cluster from ever forming.
- Request Routing: Liaison nodes look up data nodes through the registry to send query and write requests.
- Data Migration: The lifecycle agent lists all data nodes from the registry and filters them by
node_selectorlabels to pick the warm/cold targets of a migration.
BanyanDB supports three discovery mechanisms to accommodate different deployment environments:
- None Discovery: Default mode. No service discovery infrastructure is needed. Suitable for standalone deployments.
- DNS-based Discovery: Cloud-native solution leveraging Kubernetes service discovery infrastructure.
- File-based Discovery: Static configuration file approach for simple deployments and testing environments.
This document provides guidance on configuring and operating all discovery modes.
Note: As of 0.10.0, the default node discovery mode is
none. Cluster deployments must explicitly configure a discovery mode (dnsorfile). See Upgrading to 0.10 for migration notes.
Which Commands Accept These Flags
The --node-discovery-* flag set is registered in the metadata client that every long-lived BanyanDB process embeds. All four of the following commands accept the same flags and must be given a consistent configuration so that they observe the same cluster topology:
banyand databanyand liaisonbanyand standalone(onlynoneis meaningful here)lifecycle(the lifecycle agent — see the lifecycle documentation for how it uses the discovery registry to pick migration targets)
None Mode
Overview
None mode is the default discovery mode since 0.10.0. --node-discovery-mode=none disables discovery of remote peers. The discovery registry caches only the local node (if the current process carries a discoverable role) and does not locate any remote nodes. A banyand data node can still bootstrap its own property schema endpoint from its local CurNode, but liaison routing tables and lifecycle migration targets remain empty because no remote peers are ever discovered. This is the default, and it is the only valid mode for banyand standalone.
Configuration
# None mode is the default; no additional flags are required
banyand standalone
# Explicitly set none mode (equivalent to default)
banyand standalone --node-discovery-mode=none
When to Use
banyand standalone— single-process deployments that bundle the liaison and data roles together. The standalone process talks to its own in-process schema server, so no discovery is necessary.- Standalone and single-node deployments where no other nodes need to be discovered.
- Smoke tests and unit tests where a cluster is not needed and there are no external infrastructure dependencies.
Caveats
- Not valid for multi-node clusters. In
nonemode abanyand datanode can bootstrap its own property schema endpoint locally, but it will not discover any remote peers. Runningbanyand liaisonorlifecyclewith--node-discovery-mode=noneleaves liaison routing tables empty and lifecycle migrations as no-ops, because neither process can locate data nodes. Usednsorfilefor every clustered deployment. - None mode has no additional flags; there is nothing else to configure.
DNS-Based Discovery
Overview
DNS-based discovery provides a cloud-native alternative leveraging Kubernetes’ built-in service discovery infrastructure.
How it Works with Kubernetes
- Create a Kubernetes headless service.
- Kubernetes automatically creates DNS SRV records:
_<port-name>._<proto>.<service-name>.<namespace>.svc.cluster.local. - BanyanDB queries these SRV records to discover pod IPs and ports.
- Connects to each endpoint via gRPC to fetch node metadata.
Note: No external dependencies beyond DNS infrastructure.
Configuration Flags
# Mode selection
--node-discovery-mode=dns # Enable DNS mode
# DNS SRV addresses (comma-separated)
--node-discovery-dns-srv-addresses=_grpc._tcp.banyandb-data.default.svc.cluster.local
# Query intervals
--node-discovery-dns-fetch-init-interval=5s # Query interval during init (default: 5s)
--node-discovery-dns-fetch-init-duration=5m # Initialization phase duration (default: 5m)
--node-discovery-dns-fetch-interval=15s # Query interval after init (default: 15s)
# gRPC settings
--node-discovery-grpc-timeout=5s # Timeout for metadata fetch (default: 5s)
# TLS configuration
--node-discovery-dns-tls=true # Enable TLS for DNS discovery (default: false)
--node-discovery-dns-ca-certs=/path/to/ca.crt,/path/to/another.crt # Ordered CA bundle matching SRV addresses
node-discovery-dns-fetch-init-interval and node-discovery-dns-fetch-init-duration define the aggressive polling strategy during bootstrap before falling back to the steady-state
node-discovery-dns-fetch-interval. All metadata fetches share the same node-discovery-grpc-timeout, which is also reused by the file discovery mode. When TLS is enabled, the CA cert
list must match the SRV address order, ensuring each role (e.g., data, liaison) can be verified with its own trust bundle.
Configuration Examples
Single Zone Discovery:
banyand data \
--node-discovery-mode=dns \
--node-discovery-dns-srv-addresses=_grpc._tcp.banyandb-data.default.svc.cluster.local \
--node-discovery-dns-fetch-interval=10s \
--node-discovery-grpc-timeout=5s
Multi-Zone Discovery:
banyand data \
--node-discovery-mode=dns \
--node-discovery-dns-srv-addresses=_grpc._tcp.data.zone1.svc.local,_grpc._tcp.data.zone2.svc.local \
--node-discovery-dns-fetch-interval=10s
Discovery Process
The DNS discovery operates continuously in the background:
-
Query DNS SRV Records (every 5s during init, then every 15s)
- Queries each configured SRV address using
net.DefaultResolver.LookupSRV() - Deduplicates addresses across multiple SRV queries
- Queries each configured SRV address using
-
Fallback on DNS Failure
- Uses cache from previous successful query
- Allows cluster to survive temporary DNS outages
-
Update Node Cache
- For each new address: connects via gRPC, fetches node metadata, notifies handlers.
- For removed addresses: deletes from cache, notifies handlers.
Kubernetes DNS TTL
Kubernetes DNS implements Time-To-Live (TTL) caching which affects discovery responsiveness:
- Default TTL: 30 seconds (CoreDNS default since Kubernetes 1.15)
- Implications: Pod IP changes may not be visible immediately. BanyanDB polls every 15 seconds, potentially using cached results for 2-3 cycles
Customizing TTL:
To reduce discovery latency, edit the CoreDNS ConfigMap:
kubectl edit -n kube-system configmap/coredns
# Modify cache plugin TTL:
# cache 5
TLS Configuration
DNS-based discovery supports per-SRV-address TLS configuration for different roles.
Configuration Rules:
- Number of CA certificate paths must match number of SRV addresses.
caCertPaths[i]is used for connections to addresses fromsrvAddresses[i].- Multiple SRV addresses can share the same certificate file.
- If multiple SRV addresses resolve to the same IP:port, certificate from first SRV address is used.
TLS Examples:
banyand data \
--node-discovery-mode=dns \
--node-discovery-dns-tls=true \
--node-discovery-dns-srv-addresses=_grpc._tcp.data.namespace.svc.local,_grpc._tcp.liaison.namespace.svc.local \
--node-discovery-dns-ca-certs=/etc/banyandb/certs/data-ca.crt,/etc/banyandb/certs/liaison-ca.crt
Certificate Auto-Reload:
BanyanDB monitors certificate files using fsnotify and automatically reloads them when changed:
- Hot reload without process restart
- 500ms debouncing for rapid file modifications
- SHA-256 validation to detect actual content changes
This enables zero-downtime certificate rotation.
Kubernetes Deployment
Headless Service
Create a headless service for automatic DNS SRV record generation:
apiVersion: v1
kind: Service
metadata:
name: banyandb-data
namespace: default
labels:
app: banyandb
component: data
spec:
clusterIP: None # Headless service
selector:
app: banyandb
component: data
ports:
- name: grpc
port: 17912
protocol: TCP
targetPort: grpc
This creates DNS SRV record: _grpc._tcp.banyandb-data.default.svc.cluster.local
File-Based Discovery
Overview
File-based discovery provides a simple static configuration approach where nodes are defined in a YAML file. This mode is ideal for testing environments, small deployments, or scenarios where dynamic service discovery infrastructure is unavailable.
The service periodically reloads the configuration file and automatically updates the node registry when changes are detected.
How it Works
- Read node configurations from a YAML file on startup
- Attempt to connect to each node via gRPC to fetch full metadata
- Successfully connected nodes are added to the cache
- Nodes that fail to connect are queued for exponential-backoff retry by a separate retry scheduler (see the
node-discovery-file-retry-*flags below) - Reload the file at the
node-discovery-file-fetch-intervalcadence as a backup to fsnotify-based reloads; new or configuration-changed entries are processed, while unchanged failed entries continue to be retried by the backoff scheduler - Notify registered handlers when nodes are added or removed
Configuration Flags
# Mode selection
--node-discovery-mode=file # Enable file mode
# File path (required for file mode)
--node-discovery-file-path=/path/to/nodes.yaml
# gRPC settings
--node-discovery-grpc-timeout=5s # Timeout for metadata fetch (default: 5s)
# Interval settings
--node-discovery-file-fetch-interval=5m # Polling interval to reprocess the discovery file (default: 5m)
--node-discovery-file-retry-initial-interval=1s # Initial retry delay for failed node metadata fetches (default: 1s)
--node-discovery-file-retry-max-interval=2m # Upper bound for retry delay backoff (default: 2m)
--node-discovery-file-retry-multiplier=2.0 # Multiplicative factor applied between retries (default: 2.0)
node-discovery-file-fetch-interval controls the periodic full reload that acts as a safety net even if filesystem events are missed.
node-discovery-file-retry-* flags configure the exponential backoff used when a node cannot be reached over gRPC. Failed nodes are retried starting from the initial interval,
multiplied by the configured factor until the max interval is reached.
YAML Configuration Format
nodes:
- name: liaison-0
grpc_address: 192.168.1.10:17912
- name: data-hot-0
grpc_address: 192.168.1.20:17912
tls_enabled: true
ca_cert_path: /etc/banyandb/certs/ca.crt
- name: data-cold-0
grpc_address: 192.168.1.30:17912
Configuration Fields:
- name (recommended): Diagnostic label for the node, used in log messages and error output. The parser does not enforce its presence or uniqueness, but assigning a unique name to each entry is strongly recommended for operational clarity.
- grpc_address (required): gRPC endpoint in
host:portformat - tls_enabled (optional): Enable TLS for gRPC connection (default: false)
- ca_cert_path (optional): Path to CA certificate file (required when TLS is enabled)
Port Guidance: The default
--grpc-portfor bothbanyand dataandbanyand liaisonis 17912.
- Separate hosts/IPs: When data and liaison nodes run on distinct hosts or IP addresses (as shown above), both can use the default port
17912.- Co-located on one host/IP: When data and liaison share the same host or IP, they cannot bind to the same port on that IP. They also compete for the default HTTP port (
17913), observability listener (:2121), and pprof listener (:6060). Furthermore, the liaison process runs an internal pipeline server that listens on gRPC port18912and HTTP port18913by default. Overriding liaison’s public gRPC port to18912or HTTP port to18913collides with this internal listener. Therefore, when co-locating data and liaison on the same host, keep data on the defaults and configure liaison with distinct, non-conflicting public and diagnostic ports (e.g., public gRPC19912, public HTTP19913, observability:2122, and pprof:6061), pointinghttp-grpc-addrto the updated gRPC port:nodes: - name: data-0 grpc_address: 10.100.11.1:17912 - name: liaison-0 grpc_address: 10.100.11.1:19912
nodes.yamlmust list the exact gRPC port that each process actually listens on.
Configuration Examples
Basic Configuration (Data Node):
banyand data \
--node-discovery-mode=file \
--node-discovery-file-path=/etc/banyandb/nodes.yaml
Basic Configuration (Liaison Node):
# Dedicated host (default port 17912)
banyand liaison \
--node-discovery-mode=file \
--node-discovery-file-path=/etc/banyandb/nodes.yaml
# Co-located on the same host/IP as a data node:
# Assign non-conflicting public gRPC/HTTP ports (e.g. 19912/19913), redirect HTTP to the new gRPC port,
# and separate observability (:2122) and pprof (:6061) listeners.
# Note: 18912 and 18913 are used by liaison's internal pipeline server by default.
banyand liaison \
--grpc-port=19912 \
--http-port=19913 \
--http-grpc-addr=localhost:19912 \
--observability-listener-addr=:2122 \
--pprof-listener-addr=:6061 \
--node-discovery-mode=file \
--node-discovery-file-path=/etc/banyandb/nodes.yaml
With Custom Polling and Retry Settings:
banyand data \
--node-discovery-mode=file \
--node-discovery-file-path=/etc/banyandb/nodes.yaml \
--node-discovery-file-fetch-interval=20m \
--node-discovery-file-retry-initial-interval=5s \
--node-discovery-file-retry-max-interval=1m \
--node-discovery-file-retry-multiplier=1.5
Node Lifecycle
Initial/Interval Load
When the service starts:
- Reads the YAML configuration file
- Validates required fields (
grpc_address) - Attempts gRPC connection to each node
- Successfully connected nodes → added to cache
Error Handling
Startup Errors:
- Missing or invalid file path → service fails to start
- Invalid YAML format → service fails to start
- Missing required fields → service fails to start
Runtime Errors:
- gRPC connection failure → node queued for exponential-backoff retry
- File read error → keep existing cache, log error
- File deleted → keep existing cache, log error
Choosing a Discovery Mode
DNS Mode - Best For
- Kubernetes-native deployments
- Simplified operations (no external dependencies beyond DNS)
- Cloud-native architecture alignment
- StatefulSets with stable network identities
- Rapid deployment without external dependencies
File Mode - Best For
- Development and testing environments
- Small static clusters (< 10 nodes)
- Air-gapped deployments without service discovery infrastructure
- Proof-of-concept and demo setups
- Environments where node membership is manually managed
- Scenarios requiring predictable and auditable node configuration
None Mode - Best For
banyand standalone(the only supported cluster topology for this mode)- Local smoke tests and unit tests where no remote peers are involved