← all posts

Catching gRPC Schema Drift Automatically in Kubernetes

A client suddenly throws a runtime error, even though nobody changed its code. Hours of investigation often led to the same cause: someone changed a proto and deployed it to Kubernetes without updating BSR, the Buf Schema Registry. The client was built against BSR; the running server used a different schema.

I thought centralizing protos in BSR would prevent this. But a registry is a promise about what should exist, not a guarantee of what's actually running in the cluster. We thought we'd created a single source of truth. In reality, there were two truths, and they diverged.

That divergence has a name: drift. Terraform has plan, and GitOps has drift detection. Tools catch gaps between declarations and reality. I didn't have an equivalent for gRPC schemas, so I built one in Go and released it as open source.

Kubernetes made manual inspection too large a job

The check itself is simple: inspect a Pod's schema and compare it with BSR. Kubernetes complicates it. Pods span namespaces, services use different proto paths, and rolling updates mix versions within a service. More services mean linearly more repetitive checking. Repetitive work belongs in a program, not in someone's routine.

Design: extract both truths and compare them

The structure follows directly from the problem. With truth in two places, extract each and compare them.

  1. Discover gRPC Pods through the Kubernetes API.
  2. Extract live schemas with gRPC Reflection: the running reality.
  3. Download BSR schemas with the Buf CLI: the declared contract.
  4. Compare the two for mismatches.
  5. Display results in a web dashboard.

A single in-cluster Pod scans every 30 minutes.

Language: Go 1.21
Core libraries: client-go, grpcreflect, protoreflect
Architecture: Hexagonal Architecture
Deployment: In-cluster Kubernetes Pod
Storage: In-memory (sync.RWMutex)
Configuration: ConfigMap + Secret
Security: RBAC, non-root, read-only filesystem

The runtime truth lives inside the Pod

A deployed service's actual schema lives in the running binary, not its source repository. Instead of searching for proto files, I ask the running process directly.

First, client-go finds Pods. Running in-cluster lets it authenticate with the ServiceAccount token.

// Create a Kubernetes client from in-cluster configuration
config, err := rest.InClusterConfig()
clientset, err := kubernetes.NewForConfig(config)

// Load service mappings from a ConfigMap
configMap, err := clientset.CoreV1().
    ConfigMaps(namespace).
    Get(ctx, configMapName, metav1.GetOptions{})

// Find Pods by the app label
labelSelector := fmt.Sprintf("app=%s", serviceName)
pods, err := clientset.CoreV1().Pods("").
    List(ctx, metav1.ListOptions{
        LabelSelector: labelSelector,
    })

Ports are detected in order: a containerPort named grpc, another TCP port, then the default 9090.

Once a Pod is found, gRPC Reflection extracts its schema. Reflection lets a running server describe itself, exposing the runtime contract without a proto file.

// Connect to the Pod IP and port
conn, err := grpc.Dial(address, grpc.WithTransportCredentials(insecure.NewCredentials()))

// Create a Reflection client
refClient := grpcreflect.NewClientV1Alpha(ctx, reflectpb.NewServerReflectionClient(conn))

// List services
services, err := refClient.ListServices()

// Extract methods for each service
for _, serviceName := range services {
    serviceDesc, _ := refClient.ResolveService(serviceName)
}

The declared truth: replacing the HTTP API with Buf CLI

I first connected to BSR through its HTTP API and regretted it. It returned FileDescriptorSet data whose type references needed manual resolution. Parsing code kept growing, and error handling remained fragile. Rather than maintaining that complexity, I changed approaches.

// Download complete proto files with buf export
cmd := exec.CommandContext(ctx, "buf", "export", module, "-o", tmpDir)
output, err := cmd.CombinedOutput()

// Parse with protoparse
parser := protoparse.Parser{
    ImportPaths: []string{tmpDir},
}
fileDescs, err := parser.ParseFiles(relPaths...)

buf export supplies complete proto files, allowing type references to resolve and reducing errors. Installing the Buf binary in the container is a trade-off, but cheaper than maintaining my own parsing path.

Compare everything, and everything looks wrong

The first version compared all services from both sides. The dashboard turned red: test services present only at runtime were MISMATCH, and deprecated services left only in BSR were MISMATCH. Nonproblems filled the screen.

Monitoring tools can die from noise, not just missed incidents. Enough false positives teach people to ignore red lights, and then the tool may as well not exist.

I narrowed comparison to the intersection.

// Compare only services present on both sides
for liveSvcName, liveMethods := range liveServicesMap {
    if truthMethods, exists := truthServicesMap[liveSvcName]; exists {
        if !methodsMatch(liveMethods, truthMethods) {
            match = false
        }
    }
}
// Live-only services: informational; do not affect status
// BSR-only services: informational; do not affect status

Services present on only one side appear as information and don't affect status. Red indicates a service present in both places with different contents—a mismatch that can actually break a client.

Concurrency: one writer, one reader

Two goroutines run: the scanner writes every 30 minutes, and the web server reads on each request.

// 1. Scanner goroutine (every 30 minutes)
func (s *Scanner) Start(ctx context.Context) {
    ticker := time.NewTicker(scanInterval)
    for {
        select {
        case <-ticker.C:
            runScan(ctx)  // Write to the Store
        }
    }
}

// 2. Web server goroutine (per request)
func (s *Server) handleDashboard(w http.ResponseWriter, r *http.Request) {
    results := s.store.GetAll()  // Read from the Store
}

A Store protected by sync.RWMutex mediates between them. Infrequent writes and frequent reads are a straightforward fit for an RWMutex.

type Store struct {
    mu      sync.RWMutex
    results map[string]*domain.ScanResult
}

func (s *Store) GetAll() []*ScanResult {
    s.mu.RLock()
    defer s.mu.RUnlock()
    // ...
}

func (s *Store) Set(result *ScanResult) {
    s.mu.Lock()
    defer s.mu.Unlock()
    // ...
}

Architecture: replaceable adapters

I used hexagonal architecture. Kubernetes, gRPC, BSR, and the web are external adapters. The core knows only the domain operation of comparing two schemas.

protodiff/
├── cmd/protodiff/              # Entry point
├── internal/
│   ├── core/
│   │   ├── domain/             # Domain models
│   │   └── store/              # Thread-safe store
│   ├── adapters/
│   │   ├── k8s/                # Kubernetes client
│   │   ├── grpc/               # gRPC reflection client
│   │   ├── bsr/                # BSR client (Buf CLI wrapper)
│   │   └── web/                # HTTP server
│   ├── scanner/                # Orchestrator
│   └── config/                 # Configuration
└── deploy/k8s/                 # Kubernetes manifests

Go interfaces define the boundaries. Switching the BSR client from HTTP to Buf CLI validated this structure: replace and inject one adapter, and the core remains unchanged.

// BSR client interface
type Client interface {
    FetchSchema(ctx context.Context, module string) (*domain.SchemaDescriptor, error)
}

// Dependency injection
scanner := scanner.NewScanner(
    k8sClient,
    grpcClient,
    bsrClient,  // Inject through an interface
    store,
    cfg,
)

Deployment: one kubectl apply

The tool runs as a single in-cluster Pod and needs no separate infrastructure. A ConfigMap holds service mappings, a Secret holds the BSR token, and minimal RBAC permits reading Pods and ConfigMaps.

# Service mappings in a ConfigMap
apiVersion: v1
kind: ConfigMap
metadata:
  name: protodiff-mapping
  namespace: protodiff-system
data:
  user-service: "buf.build/acme/user"
  payment-service: "buf.build/acme/payment"

---
# RBAC: least privilege
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: protodiff
rules:
- apiGroups: [""]
  resources: ["pods"]
  verbs: ["get", "list", "watch"]
- apiGroups: [""]
  resources: ["configmaps"]
  verbs: ["get", "list"]
kubectl apply -f https://raw.githubusercontent.com/uzdada/protodiff/main/deploy/k8s/install.yaml

kubectl port-forward -n protodiff-system svc/protodiff 18080:80

A cluster-inspection tool can itself become an attack surface, so I included the security basics: non-root execution, a read-only filesystem, and all Linux capabilities dropped.

An observer shouldn't outweigh what it observes, so I measured its footprint: under 0.1 CPU core, roughly 10 MB of memory plus 1 KB per Pod, and a few KB of network traffic per Pod. It's I/O-bound. In-memory storage avoids an external DB. A single binary with a small operational footprint fits the MVP.

Result: a bounded detection interval

Previously, schema mismatches could remain until someone's client failed at runtime. Discovery happened through an incident, and manually comparing each service's proto could take hours.

Now, under successful scheduled scans, a mismatch is found at the next 30-minute scan. The dashboard identifies the affected service and method with SYNC, MISMATCH, or UNKNOWN. Discovering drift through an incident and discovering it before one are very different experiences.

Remaining trade-offs

The approach isn't free. Target services must enable gRPC Reflection. Go needs reflection.Register(server); Java needs server.addService(ProtoReflectionService.newInstance()), so adoption is small. In-memory history disappears on restart. Slack alerts, Prometheus metrics, PostgreSQL history, BSR caching, and CRD configuration are potential next steps.

Conclusion

Uploading a schema to a registry can feel like the job is done, but there's always room between declaration and execution. Infrastructure tools call that gap drift and measure it. ProtoDiff measures the corresponding gRPC gap every 30 minutes: discover Pods with client-go, extract runtime schemas with Reflection, fetch declarations with Buf CLI, and compare them.

It's available under Apache 2.0 for teams running gRPC microservices on Kubernetes.

GitHub: https://github.com/xhae123/protodiff