AWS 273: PrivateLink and shared services at scale
Why this lesson matters
VPC peering or Transit Gateway can connect whole routing domains when a consumer needs only one service. That broad connectivity introduces routes, overlapping-CIDR problems, and potential paths to unrelated systems. AWS PrivateLink exposes a service, resource, or controlled network segment through private endpoint addresses without making the provider network transitively routable.
The simplicity is deceptive. A production design still depends on provider load balancers or resource gateways, consumer endpoint ENIs, DNS, TLS, account permissions, acceptance, endpoint policies, application authorization, Availability Zone alignment, health, quotas, observability, cost, onboarding, and revocation. This lesson follows each layer and designs a shared service for 40 VPCs without creating hourly resources.
Outcomes
By the end, you can:
- explain PrivateLink from packet, DNS, identity, and ownership perspectives;
- distinguish interface, resource, service-network, tunnel, Gateway Load Balancer, and gateway endpoints;
- publish a custom TCP service through an NLB-backed endpoint service;
- onboard and revoke cross-account consumers safely;
- separate service permissions, connection acceptance, endpoint policy, and application authorization;
- design Regional/zonal DNS, custom private DNS, IPv4/IPv6, TLS, and source attribution;
- reason about NLB health, cross-zone behavior, AZ IDs, cross-Region access, and failure;
- compare PrivateLink with peering, TGW, VPC Lattice, API Gateway, and direct resource sharing;
- observe endpoint and provider behavior from metrics, logs, Flow Logs, and CloudTrail;
- model hourly endpoint, data-processing, load-balancer, cross-zone, DNS, and telemetry cost; and
- complete a 40-consumer service dossier with lifecycle and rollback controls.
PrivateLink mental model
An endpoint creates a narrow consumer-side entry point. For a classic custom endpoint service:
consumer application
-> service/private DNS name
-> private IP on interface endpoint ENI
-> AWS PrivateLink connection
-> provider Network Load Balancer
-> healthy target
-> application response over the established connection
The consumer sees ENIs in its own subnets and applies security groups there. The provider owns the endpoint service, NLB, listeners, target groups, targets, and service-side policy. The networks may use overlapping CIDRs because the consumer does not route directly into the provider VPC. PrivateLink is not transitive routing, and the provider cannot initiate a new connection to arbitrary consumer resources.
PrivateLink keeps service traffic on the AWS network, but it does not automatically encrypt application data. Use TLS when confidentiality, peer authentication, or application identity requires it. It also does not replace authorization: network reachability to a service is not permission to perform every operation.
Endpoint families and when to use them
| Endpoint type | Reaches | Main mechanism and boundary |
|---|---|---|
| Interface | AWS service, Marketplace service, or provider endpoint service | ENIs with private IPs in consumer subnets; usually TCP service access |
| Resource | Shared resource configuration | Direct controlled access to a database, IP, domain target, or supported resource without provider NLB |
| Service network | VPC Lattice service network | One endpoint to services/resources associated with that network |
| Tunnel | Shared network segment/CIDRs | GENEVE tunnel to a resource gateway; broader segment access needs stricter review |
| Gateway Load Balancer | Virtual appliance fleet | GENEVE traffic steering to firewalls or inspection appliances |
| Gateway | S3 and DynamoDB | Route-table prefix-list target; no endpoint ENI/hourly interface model |
Do not call every VPC endpoint an interface endpoint. Resource endpoints now allow selected resources without an NLB. Resource configurations can identify supported ARNs, domain names, private IPs, or CIDR groups and are shared through AWS RAM. Current resource application traffic is TCP; CIDR configurations use tunnel endpoints and GENEVE, with UDP limited to DNS-related behavior described by AWS.
Choose the narrowest abstraction. A single stable producer API fits an interface endpoint service. Many application services with service discovery and policy may fit VPC Lattice. General bidirectional network access fits TGW or peering. S3/DynamoDB from a VPC often fits a gateway endpoint. A database or selected IP might fit a resource endpoint. A virtual firewall fleet fits GWLB endpoints.
Provider design with endpoint services
A traditional custom endpoint service uses a Network Load Balancer. A GWLB-backed service is for virtual appliances. An NLB can belong to one endpoint service, while an endpoint service can use multiple NLBs. When multiple NLBs are present, AWS selects a same-AZ NLB for an endpoint ENI's first connection and continues using that selection, so listeners and target behavior must be consistent.
Provider sequence:
- Deploy service targets in at least two AZs and define health checks that represent useful service health.
- Create an internal NLB with listeners, target groups, supported IP type, and required AZs.
- Create an endpoint service referencing the NLB, select supported Regions and IP types, and decide whether acceptance is required.
- Grant only approved consumer principals permission to create endpoints.
- If using a friendly private DNS name, configure it and prove domain ownership.
- Publish service name, supported AZ IDs/Regions, protocol/port, DNS/TLS requirements, quotas, owner, and support contract.
- Observe pending connections and explicitly accept them when policy requires.
An NLB health check can pass while the application returns authorization errors or corrupt data. Monitor transport health and application outcomes separately. If the NLB has a security group, decide whether inbound rule evaluation for PrivateLink traffic remains enabled and define sources correctly. Never solve an attribution problem by allowing every source.
PrivateLink does not preserve consumer source IP to the target in the ordinary way. The target commonly sees load-balancer-related source information. If the application requires original connection metadata, evaluate NLB Proxy Protocol v2 and make sure the application understands it. Enabling Proxy Protocol against an unaware server corrupts the application stream. Authentication should use TLS/client identity, tokens, or signed requests rather than source IP alone.
Consumer interface endpoint design
The consumer creates an interface endpoint using the provider's service name, selects a VPC, one subnet per desired AZ, IP address type, and endpoint security groups. AWS creates an endpoint ENI in each selected subnet. Clients need DNS resolution and network reachability to those private IPs; they do not need provider CIDR routes.
Endpoint security-group inbound rules allow client connections to the service port. Client security groups need egress. NACLs must permit the flow and ephemeral return traffic. The provider's listener, target group, target security policy, TLS, and application authorization must all agree.
Use at least two supported AZs for production. Account-specific AZ names such as us-east-1a can map to different physical zones; compare Availability Zone IDs such as use1-az2 across accounts. A consumer cannot select an AZ the provider service has not enabled. Adding an AZ to an NLB does not always make it immediately usable by the endpoint service; enable it in service configuration after readiness.
PrivateLink avoids network transitivity but does not prevent a compromised workload in the consumer VPC from reaching an endpoint allowed by the same security group and DNS. Segment endpoint access by security-group reference, subnet/workload identity, endpoint policy where supported, and application authorization.
Permissions, acceptance, policy, and authorization
These controls answer different questions:
| Control | Owner | Question |
|---|---|---|
| Allowed principals | Provider | Which AWS principals may request an endpoint connection? |
| Acceptance required | Provider | Must this specific request be manually/automatically approved? |
| Endpoint policy | Consumer | Which principals/actions/resources may use a supporting AWS service through this endpoint? |
| Endpoint security group | Consumer | Which network sources and ports may reach endpoint ENIs? |
| NLB/target controls | Provider | Which connections reach healthy application targets? |
| Application identity | Provider/application | What may the authenticated caller do? |
Granting * as an allowed principal makes the service discoverable/usable more broadly according to acceptance configuration; it is not a substitute for a customer registry. Requiring acceptance without an operational queue causes onboarding delays. Automatic acceptance with broad principals can admit unintended consumers. Use specific account, role, organization, or documented patterns supported by the service design, plus event-driven review.
Endpoint policies are not universally meaningful for arbitrary provider applications. They are IAM resource policies on endpoints and their effect depends on whether the target AWS service supports them. For a custom TCP API, enforce application authentication and authorization. An endpoint policy does not override an IAM explicit deny and does not replace a service resource policy.
Revocation is layered. Remove provider permission to stop new requests, reject/delete existing connections as authorized, revoke application credentials, update DNS/service discovery, and retain audit evidence. Removing permission alone does not necessarily terminate an already accepted endpoint.
DNS, private names, and TLS
Every interface endpoint receives Regional and zonal vpce.amazonaws.com DNS names. The Regional name can return endpoint ENI addresses across enabled AZs. Zonal names target a specific AZ and can support locality, but the application must handle zonal failure.
For integrated AWS services, private DNS can create an AWS-managed hidden private hosted zone so the normal public service hostname resolves to endpoint private IPs inside the VPC. For a provider endpoint service, the provider can associate a private DNS name after proving domain ownership; consumers then enable that behavior where supported. Alternatively, consumers create Route 53 private records pointing to endpoint aliases.
Check VPC DNS support/hostnames, private-zone conflicts, Resolver rules, Profiles, on-premises forwarding, and split-view behavior from AWS272. Endpoint-generated DNS names are publicly resolvable but return private addresses, which remain unreachable without private network access to the consumer VPC.
TLS names must match the hostname clients use. A certificate for api.internal.example will not validate an endpoint-specific vpce-... name. Decide where TLS terminates, how SNI reaches the right listener/target, who rotates certificates, and whether mutual TLS is required. DNS failover cannot repair an unhealthy application if all aliases still target the same broken service.
IPv4, IPv6, dual-stack endpoint types, and DNS record IP types must be compatible with the endpoint service and clients. Do not assume that dual-stack consumer subnets make an IPv4-only provider service support IPv6.
Zonal behavior, scaling, and failure
PrivateLink is Regional but endpoint ENIs are zonal. Prefer same-AZ client-to-endpoint paths to reduce cross-AZ dependencies/cost, while retaining another AZ for failure. The provider should place healthy targets in each offered AZ. If targets are absent in an enabled AZ, cross-zone NLB load balancing can reach another AZ but adds dependency and may add Regional data-transfer charges.
Test:
- one endpoint ENI/security group/path failure;
- one provider target and full target-group failure;
- one provider NLB AZ failure;
- DNS cache behavior during endpoint replacement;
- connection draining and long-lived TCP sessions;
- consumer retry/backoff and idempotency;
- quota exhaustion during mass onboarding; and
- provider deployment rollback while existing connections continue.
PrivateLink scales the managed connection layer, not the provider application. NLB capacity, target limits, connection rates, ephemeral ports, TLS handshakes, application pools, database connections, and downstream dependencies still need load tests and alarms.
Cross-Region access can expose an endpoint service from its service Region to supported consumer Regions selected by the provider. Treat it as an explicit architecture choice: inspect supported Regions, principal permissions, DNS, data-transfer charges, latency, failure ownership, and the documented metrics limitation for cross-Region consumers. It is not a multi-Region active-active application by itself.
Shared-service operating model for 40 VPCs
Use a service catalog entry containing owner, data classification, service name, private DNS name, ports/protocols, supported Regions/AZ IDs/IP types, availability target, authentication, endpoint-policy support, quota, cost allocation, monitoring, support, onboarding, revocation, and deprecation date.
Avoid 40 hand-built endpoints. Use reviewed infrastructure as code, account vending hooks, AWS Organizations metadata, and a consumer registry. Validate each request against account, VPC, environment, Region, AZ IDs, subnet capacity, security group, DNS conflict, application entitlement, and cost center. Tag endpoint and connection records consistently.
Decide between distributed and centralized endpoints:
- One endpoint per consumer VPC preserves isolation, simple DNS, and clear chargeback but increases endpoint-hour cost and inventory.
- Central endpoint VPC accessed through TGW/peering may reduce endpoint count but reintroduces routing, transitivity controls, centralized DNS, cross-AZ data, blast radius, and chargeback complexity.
PrivateLink's isolation advantage is weakened when consumers route through a shared endpoint VPC. Select centralization only after modeling the complete path and failure domain.
Read-only evidence collection
In VPC > Endpoint services, inspect service name/state, NLB/GWLB, enabled AZs/Regions, supported IP types, acceptance, private DNS verification, allowed principals, and endpoint connections. In VPC > Endpoints, inspect type, service/resource, state, VPC/subnets/ENIs, DNS, policy, route tables where applicable, and security groups.
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ec2 describe-vpc-endpoint-service-configurations --output json
aws ec2 describe-vpc-endpoint-connections --output json
aws ec2 describe-vpc-endpoints --output json
aws ec2 describe-vpc-endpoint-services --output json
For approved identifiers:
SERVICE_ID="vpce-svc-approved-id"
ENDPOINT_ID="vpce-approved-id"
aws ec2 describe-vpc-endpoint-service-permissions + --service-id "$SERVICE_ID" --output json
aws ec2 describe-vpc-endpoint-service-configurations + --service-ids "$SERVICE_ID" --output json
aws ec2 describe-vpc-endpoints + --vpc-endpoint-ids "$ENDPOINT_ID" --output json
Also collect NLB listeners, target health, attributes, endpoint ENIs, security groups, DNS records, CloudWatch metrics, VPC Flow Logs, access/application logs, CloudTrail, Service Quotas, RAM resource configurations, and Cost and Usage Report dimensions. Redact account IDs, service names, private domains, addresses, and topology.
Monitoring and diagnosis
PrivateLink publishes endpoint and endpoint-service metrics including active connections, new connections, bytes, and rejected/reset behavior where applicable. NLB metrics expose target health, flows, resets, TLS, and capacity. Flow Logs prove ENI-level accepts/rejects but not TLS or application success.
| Symptom | First evidence | Likely fault |
|---|---|---|
| Service name not found/unauthorized | service permission and Region | wrong principal, Region, or service name |
| Endpoint pending acceptance | provider connection queue | acceptance workflow not completed |
| Endpoint rejected/failed | connection state/status message | provider rejection or incompatible configuration |
| DNS returns public address | private DNS/PHZ/Resolver path | DNS disabled, conflict, wrong query origin |
| DNS returns private IP, timeout | endpoint SG/Flow Logs/path | port/source/NACL/client egress |
| TCP connects, TLS fails | SNI/certificate chain/time | wrong hostname, trust, expired certificate |
| 401/403 after TLS | application/IAM logs | network works; identity or authorization denied |
| Intermittent by AZ | zonal DNS, ENI and target health | unsupported/unhealthy provider AZ |
| Provider sees unexpected source | NLB/Proxy Protocol design | source-preservation assumption |
| Existing consumer works after revocation | connection/application inventory | permission removal stopped only new endpoints |
Diagnose in order: exact hostname and resolved IP, endpoint state, connection acceptance, client-to-ENI security, service/AZ compatibility, listener and target health, TLS, application identity, downstream dependency. Never widen principals, endpoint security groups, or target rules globally to test.
Decision boundaries
Use PrivateLink when consumers need narrow private access, provider/consumer CIDRs may overlap, the provider must hide its topology, and one-way connection initiation is acceptable. It is strong for internal APIs, SaaS, AWS services, and appliance insertion.
Reject or combine it when clients need broad bidirectional network connectivity, unsupported protocols, multicast, direct host identity by source IP, consumer-initiated routing to many changing ports, or transitive network access. Compare:
| Requirement | Better starting choice |
|---|---|
| One provider TCP service to many accounts | NLB endpoint service |
| Selected database/IP/domain target | Resource endpoint/configuration |
| Many discoverable services with auth/routing | VPC Lattice |
| Whole-network routed connectivity | TGW or peering |
| S3/DynamoDB private access | Gateway endpoint |
| Public/mobile API with edge controls | API Gateway/ALB design |
| Transparent virtual appliance insertion | GWLB endpoint service |
Forty-consumer architecture workbook
Design a provider API consumed by 20 production and 20 development VPCs across four accounts and two Regions. Production requires TLS, explicit approval, two-AZ availability, chargeback, and 30-day revocation. Development has lower capacity but cannot call production operations.
For every consumer, record:
| Field | Required proof |
|---|---|
| Identity | account/principal, environment, application owner |
| Placement | Region, AZ IDs, subnets, address/IP type |
| Connection | endpoint ID/state, provider acceptance, service version |
| Network | endpoint SG, client SG, NACL, port/protocol |
| DNS/TLS | Regional/zonal/custom name, PHZ conflict, SNI/certificate |
| Authorization | endpoint-policy support, application role/scope |
| Resilience | second AZ, retry, target health, failure result |
| Observability | endpoint/NLB metrics, Flow/app/audit logs |
| Cost | endpoint hours, data, NLB, transfer, logs, cost center |
| Lifecycle | onboarding, owner review, revoke, replace, delete evidence |
Submit the 40 rows, endpoint-type decision, provider/consumer diagrams, AZ-ID map, permissions/acceptance matrix, DNS/TLS plan, quota forecast, cost model, alarms, five failure tests, revocation drill, service replacement plan, and RACI. Include proof that no resource changed.
Cost, quotas, change, and rollback
Consumers generally pay interface endpoint hours per AZ and tiered data processing; provider, cross-Region, resource, service-network, GWLB, and tunnel models have their own current pricing dimensions. Add NLB/GWLB hours and capacity, targets, cross-zone or cross-Region transfer, Route 53, logs, NAT/TGW if introduced, and support operations. Forty VPCs times two or three AZs creates many billable endpoint-hours even at zero traffic.
Quotas cover endpoints per VPC/Region, services, connections, ENIs, subnets/AZs, principals, load balancer targets/listeners, resource configurations, and bandwidth/flow characteristics. Retrieve current Service Quotas in the deployment Region instead of copying a static number into a design.
Before changes, snapshot service configuration, NLBs/AZs, permissions, connections, endpoints, policies, SGs, DNS, certificates, target health, alarms, tags, and costs. Canary one consumer, hold telemetry gates, and preserve the previous service/DNS during replacement. Rollback may restore permission/acceptance, endpoint subnets/SG, old DNS alias, NLB/target group, certificate, or application version. Existing DNS caches and long-lived connections affect recovery time.
This lesson creates nothing. Before/after evidence must show identical endpoint services, permissions, connections, endpoints, NLBs, resource configurations, DNS, SGs, and routes.
Knowledge check
- Does PrivateLink route the consumer CIDR into the provider VPC?
No. The consumer connects to endpoint ENIs and the managed service path reaches the provider service.
- Why compare AZ IDs across accounts?
AZ letters can map differently, while AZ IDs identify the same physical zone.
- Does acceptance authorize API operations?
No. It approves the endpoint connection; application/IAM authorization remains separate.
- Does PrivateLink encrypt the application payload?
Not by itself. Use TLS where required.
- Why can removing allowed-principal permission be incomplete revocation?
It prevents new use according to configuration but an existing accepted connection and credentials may remain.
- When is a resource endpoint preferable?
When a supported selected resource/IP/domain should be shared directly without an NLB service.
Lesson acceptance
The lesson is complete only when the learner can:
- trace consumer DNS, endpoint ENI, managed connection, NLB/resource gateway, target, and response;
- select the correct endpoint family and reject PrivateLink for routed-network requirements;
- design provider and consumer controls across 40 VPCs;
- separate permissions, acceptance, endpoint policy, security groups, TLS, and application authorization;
- prove AZ-ID alignment, two-AZ health, DNS/TLS, IPv4/IPv6, and cross-Region behavior;
- diagnose control-plane, network, TLS, identity, and application failures from evidence;
- model quotas, all recurring cost owners, onboarding, revocation, replacement, and rollback; and
- attest that no staging or production AWS resource changed.