[CloudNeta] Hands-On LLM Serving 6주차 part 1 - Scaling LLM Inference with vLLM and AWS Tranium Workshop

연재 안내

이 글은 6주차 연재의 첫 번째 글입니다.

  1. part 1 - Scaling LLM Inference with vLLM and AWS Tranium Workshop (현재 글)

들어가며

그림 1 - Test Client의 요청이 ELB와 NGINX Ingress를 거쳐 Neuron 기반 vLLM 파드로 전달되고, S3 캐시와 AWS 관리형 모니터링 서비스로 연결되는 실습 환경

Lab 1: EKS 클러스터 설정

사전준비

아래 명령어로 환경구성을 진행합니다.

#
cd workshop
pwd

# Update package list and install tools
#   커널 재부팅이 있으니 진행합시다. 바로 진행되네요
echo "Updating package list and installing tools..."
sudo apt update
sudo apt install -y python3-pip jq unzip

# Install AWS CLI v2
#   별 문제없이 잘 됩니다
echo "Installing AWS CLI v2..."
curl "https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip" -o "awscliv2.zip"
unzip awscliv2.zip
sudo ./aws/install --update

# Install Helm
#   헬름 설치도 잘 되고요
echo "Installing Helm..."
curl https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash

# Enable kubectl autocompletion for current session and add to bashrc
echo "Setting up kubectl autocompletion..."
source <(kubectl completion bash) && echo "source <(kubectl completion bash)" >> ~/.bashrc

# Verify installations
echo "Verifying installations..."
aws --version
helm version --short
jq --version
echo "kubectl autocompletion enabled!"

이런 식으로 잘 되는 것을 볼 수 있습니다.

ubuntu@ip-10-0-1-41:~/workshop$ echo "Verifying installations..."
aws --version
helm version --short
jq --version
echo "kubectl autocompletion enabled!"
Verifying installations...
aws-cli/2.36.44 Python/3.14.6 Linux/6.8.0-1035-aws exe/x86_64.ubuntu.22
v3.22.0+g144ca65
jq-1.6
kubectl autocompletion enabled!

이어서 환경변수 세팅을 합시다. 꼭 해주고 넘어가야하니 참고해주세요

export AWS_REGION=us-west-2
export CLUSTER_NAME=ai-infra-summit-test-cluster
export EKS_VERSION=1.33
export INSTANCE_TYPE=trn1.2xlarge
export DESIRED_NODES=1
export WORKER_AMI=$(aws ssm get-parameter \
    --name /aws/service/eks/optimized-ami/1.33/amazon-linux-2023/x86_64/neuron/recommended/image_id \
    --region $AWS_REGION \
    --query "Parameter.Value" \
    --output text)
export AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
export BUCKET_NAME=ai-infra-summit-vllm-models-cache-${AWS_ACCOUNT_ID}

echo "$CLUSTER_NAME $WORKER_AMI $AWS_ACCOUNT_ID $BUCKET_NAME"

잘 되네요!

echo "$CLUSTER_NAME $WORKER_AMI $AWS_ACCOUNT_ID $BUCKET_NAME"
ai-infra-summit-test-cluster ami-REDACTED REDACTED ai-infra-summit-vllm-models-cache-REDACTED

클러스터 접근 설정

aws eks update-kubeconfig 으로 접근권한을 얻은 후 파드와 클러스터 정보를 확인해봅시다.

$ ubuntu@ip-10-0-1-41:~/workshop$ aws eks update-kubeconfig --region $AWS_REGION --name $CLUSTER_NAME
Added new context arn:aws:eks:us-west-2:<REDACTED>:cluster/ai-infra-summit-test-cluster to /home/ubuntu/.kube/config

$ ubuntu@ip-10-0-1-41:~/workshop$ cat ~/.kube/config
apiVersion: v1
clusters:
- cluster:
    certificate-authority-data: <REDACTED>
    server: https://D5758DECBE4C8436A000042718FC282D.gr7.us-west-2.eks.amazonaws.com
  name: arn:aws:eks:us-west-2:<REDACTED>:cluster/ai-infra-summit-test-cluster
contexts:
- context:
    cluster: arn:aws:eks:us-west-2:<REDACTED>:cluster/ai-infra-summit-test-cluster
    user: arn:aws:eks:us-west-2:<REDACTED>:cluster/ai-infra-summit-test-cluster
  name: arn:aws:eks:us-west-2:<REDACTED>:cluster/ai-infra-summit-test-cluster
current-context: arn:aws:eks:us-west-2:<REDACTED>:cluster/ai-infra-summit-test-cluster
kind: Config
preferences: {}
users:
- name: arn:aws:eks:us-west-2:<REDACTED>:cluster/ai-infra-summit-test-cluster
  user:
    exec:
      apiVersion: client.authentication.k8s.io/v1beta1
      args:
      - --region
      - us-west-2
      - eks
      - get-token
      - --cluster-name
      - ai-infra-summit-test-cluster
      - --output
      - json
      command: aws

현황을 봅시다.

$ ubuntu@ip-10-0-1-41:~/workshop$ kubectl get pod -A
NAMESPACE     NAME                       READY   STATUS    RESTARTS   AGE
kube-system   coredns-75cb89d95b-gmbhv   0/1     Pending   0          37h
kube-system   coredns-75cb89d95b-tht4q   0/1     Pending   0          37h

$ ubuntu@ip-10-0-1-41:~/workshop$ kubectl cluster-info
Kubernetes control plane is running at https://REDACTED.gr7.us-west-2.eks.amazonaws.com
CoreDNS is running at https://REDACTED.gr7.us-west-2.eks.amazonaws.com/api/v1/namespaces/kube-system/services/kube-dns:dns/proxy

To further debug and diagnose cluster problems, use 'kubectl cluster-info dump'.

이후 k9s 를 설치합니다.

그림 2 - k9s에서 확인한 Kubernetes 클러스터 리소스

클러스터 네트워킹 설정

이후 노드 접근작업의 편의를 위해 ssh 키를 추가하고, 클러스터 네트워킹 설정을 합시다.

# Get VPC and create public subnets for trn1.2xlarge instances
VPC_ID=$(aws eks describe-cluster --name $CLUSTER_NAME --region $AWS_REGION --query 'cluster.resourcesVpcConfig.vpcId' --output text)
PUBLIC_ROUTE_TABLE=$(aws ec2 describe-route-tables --filters "Name=vpc-id,Values=$VPC_ID" "Name=route.destination-cidr-block,Values=0.0.0.0/0" --query 'RouteTables[0].RouteTableId' --output text)
echo "$VPC_ID $PUBLIC_ROUTE_TABLE"
vpc-06cda06a1da780126 rtb-0d694588705be058d
# Get supported AZs for instance type
SUPPORTED_AZS=($(aws ec2 describe-instance-type-offerings --location-type availability-zone --filters "Name=instance-type,Values=$INSTANCE_TYPE" --query 'InstanceTypeOfferings[*].Location' --output text))
echo "SUPPORTED_AZS" $SUPPORTED_AZS
SUPPORTED_AZS us-west-2b
# Get public subnets in supported AZs
VALID_SUBNETS=()
for az in "${SUPPORTED_AZS[@]}"; do
  subnet=$(aws ec2 describe-subnets --filters "Name=vpc-id,Values=$VPC_ID" "Name=map-public-ip-on-launch,Values=true" "Name=availability-zone,Values=$az" --query 'Subnets[0].SubnetId' --output text)
  [ "$subnet" != "None" ] && [ "$subnet" != "" ] && VALID_SUBNETS+=("$subnet")
done
echo ${VALID_SUBNETS[0]}
echo ${VALID_SUBNETS[1]}
subnet-028a53e0105798217
subnet-01e93e746cefe4850
# Ensure we have at least 2 subnets
[ ${#VALID_SUBNETS[@]} -lt 2 ] && { echo "Error: Need at least 2 public subnets in AZs that support $INSTANCE_TYPE"; exit 1; }
#   에러 없이 아무것도 안 나오니까 잘 됐습니다.
# Get first two valid subnets and their AZs
PUBLIC_SUBNET_1=${VALID_SUBNETS[0]}
PUBLIC_SUBNET_2=${VALID_SUBNETS[1]}
AZ_1=$(aws ec2 describe-subnets --subnet-ids $PUBLIC_SUBNET_1 --query 'Subnets[0].AvailabilityZone' --output text)
AZ_2=$(aws ec2 describe-subnets --subnet-ids $PUBLIC_SUBNET_2 --query 'Subnets[0].AvailabilityZone' --output text)
echo $AZ_1 $AZ_2
us-west-2b us-west-2d

성공적으로 나오는 걸 확인할 수 있습니다.

echo "Using PUBLIC_SUBNET_1: $PUBLIC_SUBNET_1 in $AZ_1"
echo "Using PUBLIC_SUBNET_2: $PUBLIC_SUBNET_2 in $AZ_2"
Using PUBLIC_SUBNET_1: subnet-028a53e0105798217 in us-west-2b
Using PUBLIC_SUBNET_2: subnet-01e93e746cefe4850 in us-west-2d

노드그룹 배포

이어서 아래 파일을 생성하고 노드그룹을 배포해봅시다.

ubuntu@ip-10-0-1-41:~/workshop$ eksctl create nodegroup --config-file=eks_nodegroup.yaml
2026-09-12 18:30:40 [!]  no eksctl-managed CloudFormation stacks found for "ai-infra-summit-test-cluster", will attempt to create nodegroup(s) on non eksctl-managed cluster
2026-09-12 18:30:40 [ℹ]  nodegroup "neuron-trn1-2x" will use "ami-0e08c07b0376ba3f8" [AmazonLinux2023/1.33]
2026-09-12 18:30:41 [ℹ]  using SSH public key "/home/ubuntu/.ssh/id_rsa.pub" as "eksctl-ai-infra-summit-test-cluster-nodegroup-neuron-trn1-2x-4e:6c:aa:e7:97:66:5b:1b:d9:b3:5f:c9:c8:cf:3b:79" 
2026-09-12 18:30:41 [ℹ]  1 nodegroup (neuron-trn1-2x) was included (based on the include/exclude rules)
2026-09-12 18:30:41 [ℹ]  will create a CloudFormation stack for each of 1 managed nodegroups in cluster "ai-infra-summit-test-cluster"
2026-09-12 18:30:41 [ℹ]  1 task: { 1 task: { 1 task: { create managed nodegroup "neuron-trn1-2x" } } }
2026-09-12 18:30:41 [ℹ]  building managed nodegroup stack "eksctl-ai-infra-summit-test-cluster-nodegroup-neuron-trn1-2x"
2026-09-12 18:30:41 [ℹ]  deploying stack "eksctl-ai-infra-summit-test-cluster-nodegroup-neuron-trn1-2x"
2026-09-12 18:30:41 [ℹ]  waiting for CloudFormation stack "eksctl-ai-infra-summit-test-cluster-nodegroup-neuron-trn1-2x"

...

잘 떴으니 이를 확인하고

ubuntu@ip-10-0-1-41:~/workshop$ k get nodes -o wide
NAME                                      STATUS   ROLES    AGE   VERSION                INTERNAL-IP   EXTERNAL-IP     OS-IMAGE            KERNEL-VERSION                     CONTAINER-RUNTIME
ip-10-0-2-86.us-west-2.compute.internal   Ready    <none>   14m   v1.33.13-eks-cb19647   10.0.2.86     44.243.86.194   Amazon Linux 2023.12.20260831   6.12.103-127.188.amzn2023.x86_64   containerd://2.2.5+unknown

neuron 디바이스 플러그인을 확인합니다.

ubuntu@ip-10-0-1-41:~/workshop$ k get pods -n kube-system | grep neuron
neuron-device-plugin-tb42j   1/1     Running   0          13m
k describe pod -n kube-system -l name=neuron-device-plugin-ds
Name:                 neuron-device-plugin-tb42j
Namespace:            kube-system
Priority:             2000001000
Priority Class Name:  system-node-critical
Service Account:      neuron-device-plugin
Node:                 ip-10-0-2-86.us-west-2.compute.internal/10.0.2.86
Start Time:           Sat, 12 Sep 2026 18:33:09 +0000
Labels:               app.kubernetes.io/name=neuron-device-plugin
                      controller-revision-hash=6fcd58bb84
                      name=neuron-device-plugin-ds
                      pod-template-generation=1
Annotations:          <none>
Status:               Running
IP:                   10.0.2.89
IPs:
  IP:           10.0.2.89
Controlled By:  DaemonSet/neuron-device-plugin
Containers:
  neuron-device-plugin:
    Container ID:   containerd://da74a9cb5304184d59c0d1e3e5adbf760db90d64e418b29b596ccd54ca005470
    Image:          public.ecr.aws/neuron/neuron-device-plugin:2.23.30.0
    Image ID:       public.ecr.aws/neuron/neuron-device-plugin@sha256:75a6d5ce3bd397c4d05ce7dd4b51a306c7f0a0e1c146710029d5144caed94aa1
    Port:           <none>
    Host Port:      <none>

... 이하 생략
ubuntu@ip-10-0-1-41:~/workshop$ k exec -it -n kube-system ds/neuron-device-plugin -- ls -l /var/lib/kubelet/device-plugins
total 4
srwxr-xr-x. 1 root root   0 Sep 12 18:32 kubelet.sock
-rw-------. 1 root root 145 Sep 12 18:33 kubelet_internal_checkpoint
srwxr-xr-x. 1 root root   0 Sep 12 18:33 neuron-devplugin.sock
srwxr-xr-x. 1 root root   0 Sep 12 18:33 neuroncore-devplugin.sock
ubuntu@ip-10-0-1-41:~/workshop$ k exec -it -n kube-system ds/neuron-device-plugin -- ls -l /opt/aws
total 16
drwxr-xr-x. 3 root root    45 Sep  3 03:18 apitools
drwxr-xr-x. 2 root root 16384 Sep  3 03:18 bin
drwxr-xr-x. 5 root root    41 Sep  3 03:22 neuron
ubuntu@ip-10-0-1-41:~/workshop$ k exec -it -n kube-system ds/neuron-device-plugin -- ls -R -1 /opt/aws/neuron
/opt/aws/neuron:
bin
lib
share

/opt/aws/neuron/bin:
api_pb2.py
api_pb2_grpc.py
default-slurm-setup.sh
nccom-test
neuron-bench
neuron-dbg
neuron-dump
neuron-dump.py
neuron-explorer
neuron-ls
neuron-monitor
neuron-monitor-cloudwatch.py
neuron-monitor-device-view.py
neuron-monitor-k8s-info.py
neuron-monitor-prometheus.py
neuron-monitor-top.py
neuron-profile
neuron-top

/opt/aws/neuron/lib:
libndbg.so

/opt/aws/neuron/share:
man

/opt/aws/neuron/share/man:
man1

/opt/aws/neuron/share/man/man1:
neuron-ls.1
neuron-monitor.

이렇게 하면 새 노드가 잘 떠있는 것을 확인할 수 있습니다.

그림 3 - 새로 생성한 EKS 노드와 클러스터 상태

직접 session manager로 들어가서 확인해보죠.

그림 4 - Neuron 장치와 코어, 메모리 정보

neuron cache용 버킷 생성

neuron-cache 를 담을 버킷을 생성합니다.

ubuntu@ip-10-0-1-41:~/workshop$ echo $BUCKET_NAME
ai-infra-summit-vllm-models-cache-908254650436
ubuntu@ip-10-0-1-41:~/workshop$ aws s3 mb "s3://$BUCKET_NAME" --region "$AWS_REGION"
make_bucket: ai-infra-summit-vllm-models-cache-908254650436
ubuntu@ip-10-0-1-41:~/workshop$ aws s3 ls
2026-09-12 18:58:30 ai-infra-summit-vllm-models-cache-908254650436

neuron device plugin 재설치

neuron device plugin을 재설치하고 neuron 스케줄러 확장을 설치합니다.

재설치 이유?

neuron-device-plugin과 향후 Helm으로 설치할 플러그인이 충돌하기 때문입니다.
사전에 설치한 리소스는 Helm으로 관리되지 않아서 k8s의 관리하에 들어오지 않아 이를 싱크맞추기 위함입니다.

그러므로 싹 지우고,

# Clean Up Any Existing Neuron Components
k delete daemonset neuron-device-plugin -n kube-system
k delete clusterrole neuron-device-plugin
k delete serviceaccount neuron-device-plugin -n kube-system
k delete clusterrolebinding neuron-device-plugin

daemonset.apps "neuron-device-plugin" deleted from kube-system namespace
clusterrole.rbac.authorization.k8s.io "neuron-device-plugin" deleted
serviceaccount "neuron-device-plugin" deleted from kube-system namespace
clusterrolebinding.rbac.authorization.k8s.io "neuron-device-plugin" deleted

Helm으로 아무것도 없는걸 확인하고, 마저 설치합시다.

ubuntu@ip-10-0-1-41:~/workshop$ helm list -A
NAME    NAMESPACE       REVISION        UPDATED STATUS  CHART   APP VERSION

ubuntu@ip-10-0-1-41:~/workshop$ helm upgrade --install neuron-helm-chart oci://public.ecr.aws/neuron/neuron-helm-chart --set "npd.enabled=false"
Release "neuron-helm-chart" does not exist. Installing it now.
Pulled: public.ecr.aws/neuron/neuron-helm-chart:1.10.0
Digest: sha256:ed5d8f73b7a05d3a1edf17b2bfaf277ea995ad1c65b115fa8ba2c7b4dc649346
NAME: neuron-helm-chart
LAST DEPLOYED: Sat Sep 12 19:04:18 2026
NAMESPACE: default
STATUS: deployed
REVISION: 1
NOTES:

ubuntu@ip-10-0-1-41:~/workshop$ helm list -A
NAME                    NAMESPACE       REVISION        UPDATED                                 STATUS          CHART                     APP VERSION
neuron-helm-chart       default         1               2026-09-12 19:04:18.94730929 +0000 UTC  deployed        neuron-helm-chart-1.10.0  1.10.0 

새로뜬걸 확인하고, neuron core와 디바이스가 잘 할당되었나 봅시다.

ubuntu@ip-10-0-1-41:~/workshop$ k get ds neuron-device-plugin -n kube-system
NAME                   DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   NODE SELECTOR   AGE
neuron-device-plugin   1         1         1       1            1           <none>          65s

ubuntu@ip-10-0-1-41:~/workshop$ k get nodes "-o=custom-columns=NAME:.metadata.name,NeuronCore:.status.allocatable.aws\.amazon\.com/neuroncore"
NAME                                      NeuronCore
ip-10-0-2-86.us-west-2.compute.internal   2

이후 neuron-scheduler-extension을 설치합시다.

ubuntu@ip-10-0-1-41:~/workshop$ helm upgrade --install neuron-helm-chart oci://public.ecr.aws/neuron/neuron-helm-chart \
    --set "scheduler.enabled=true" \
    --set "npd.enabled=false"
Pulled: public.ecr.aws/neuron/neuron-helm-chart:1.10.0
Digest: sha256:ed5d8f73b7a05d3a1edf17b2bfaf277ea995ad1c65b115fa8ba2c7b4dc649346
Release "neuron-helm-chart" has been upgraded. Happy Helming!
NAME: neuron-helm-chart
LAST DEPLOYED: Sat Sep 12 19:06:38 2026
NAMESPACE: default
STATUS: deployed
REVISION: 2
NOTES:

잘 떠있는지 확인해볼까요.

ubuntu@ip-10-0-1-41:~/workshop$ k get pod -n kube-system -l app.kubernetes.io/component=k8s-neuron-scheduler
k get deploy -n kube-system k8s-neuron-scheduler

NAME                                    READY   STATUS    RESTARTS   AGE
k8s-neuron-scheduler-785c8d99f8-n2lt8   1/1     Running   0          27s
NAME                   READY   UP-TO-DATE   AVAILABLE   AGE
k8s-neuron-scheduler   1/1     1            1           28s



ubuntu@ip-10-0-1-41:~/workshop$ k get pod -n kube-system -l app.kubernetes.io/component=my-scheduler
k get deploy -n kube-system my-scheduler
NAME                            READY   STATUS    RESTARTS   AGE
my-scheduler-55f56bc9f8-8k42m   1/1     Running   0          61s
NAME           READY   UP-TO-DATE   AVAILABLE   AGE
my-scheduler   1/1     1            1           62s

Amazon S3 CSI 드라이버 설치하기

쿠버네티스에서 S3에 붙을 수 있도록 CSI 드라이버를 설치합시다.

ubuntu@ip-10-0-1-41:~/workshop$ helm repo add aws-mountpoint-s3-csi-driver https://awslabs.github.io/mountpoint-s3-csi-driver
"aws-mountpoint-s3-csi-driver" has been added to your repositories
ubuntu@ip-10-0-1-41:~/workshop$ helm repo update
Hang tight while we grab the latest from your chart repositories...
...Successfully got an update from the "aws-mountpoint-s3-csi-driver" chart repository
Update Complete. ⎈Happy Helming!⎈
ubuntu@ip-10-0-1-41:~/workshop$ helm upgrade --install aws-mountpoint-s3-csi-driver \
    --namespace kube-system \
    aws-mountpoint-s3-csi-driver/aws-mountpoint-s3-csi-driver
Release "aws-mountpoint-s3-csi-driver" does not exist. Installing it now.
NAME: aws-mountpoint-s3-csi-driver
LAST DEPLOYED: Sat Sep 12 19:08:42 2026
NAMESPACE: kube-system
STATUS: deployed
REVISION: 1
TEST SUITE: None
NOTES:
Thank you for using Mountpoint for Amazon S3 CSI Driver v2.8.0.

Learn more about the file system operations Mountpoint supports: https://github.com/awslabs/mountpoint-s3/blob/main/doc/SEMANTICS.md

잘 떠있군요!

ubuntu@ip-10-0-1-41:~/workshop$ kubectl get pods -n kube-system -l app.kubernetes.io/name=aws-mountpoint-s3-csi-driver
NAME                                 READY   STATUS    RESTARTS   AGE
s3-csi-controller-5df587766f-s8vz2   1/1     Running   0          21s
s3-csi-node-gnjmm                    3/3     Running   0          21s

클러스터가 잘 살아있나 살펴봅시다.

ubuntu@ip-10-0-1-41:~/workshop$ echo "=== Cluster Status ==="
kubectl get nodes

echo -e "\n=== Neuron Devices ==="
kubectl describe nodes -l alpha.eksctl.io/nodegroup-name=neuron-trn1-2x | grep "aws.amazon.com/neuron"

echo -e "\n=== Storage Classes ==="
kubectl get storageclass

echo -e "\n=== Current Namespace ==="
kubectl config get-contexts
=== Cluster Status ===
NAME                                      STATUS   ROLES    AGE   VERSION
ip-10-0-2-86.us-west-2.compute.internal   Ready    <none>   41m   v1.33.13-eks-cb19647

=== Neuron Devices ===
  aws.amazon.com/neuron:      1
  aws.amazon.com/neuroncore:  2
  aws.amazon.com/neuron:      1
  aws.amazon.com/neuroncore:  2
  aws.amazon.com/neuron      0           0
  aws.amazon.com/neuroncore  0           0

=== Storage Classes ===
NAME   PROVISIONER             RECLAIMPOLICY   VOLUMEBINDINGMODE      ALLOWVOLUMEEXPANSION   AGE
gp2    kubernetes.io/aws-ebs   Delete          WaitForFirstConsumer   false                  38h

=== Current Namespace ===
CURRENT   NAME                                                                      CLUSTER                   AUTHINFO                                                                  NAMESPACE
*         arn:aws:eks:us-west-2:REDACTED:cluster/ai-infra-summit-test-cluster   arn:aws:eks:us-west-2:REDACTED:cluster/ai-infra-summit-test-cluster   arn:aws:eks:us-west-2:REDACTED:cluster/ai-infra-summit-test-cluster

지금까지 아래 내용을 구성했습니다.

추가: NVIDIA GPU 리소스 vs Neuron GPU 리소스

보통은 그럼 컨테이너/파드가 NVIDIA GPU를 쓸텐데, 이럼 Neuron GPU 리소스와는 어떤 차이가 있을까요?

컨테이너 런타임 레벨 vs 순수 k8s 디바이스 플러그인 레벨 개입

NVIDIA 설정

그림 5 - NVIDIA GPU 리소스가 파드에 노출되는 방식

그림 6 - NVIDIA CDI 설정이 컨테이너 런타임에 포함되는 방식

Neuron 설정

통상의 엔비디아 GPU 리소스 활용의 경우 컨테이너 런타임이 시스템 깊숙하게 관여해야하기 때문이고, Neuron GPU 리소스는 커널레벨, 유저영역의 SDK 관리를 해두었기 때문에 가능합니다. Neuron GPU의 세밀한 제어는 전용 커스텀 스케줄러 확장이 필요합니다.

Lab 2: vLLM 배포

여기서는 초기화 컨테이너 패턴을 쓰는 파드를 통해 EKS 클러스터에 vLLM을 배포해볼 예정입니다.

허깅페이스 토큰 관리용 시크릿 생성

허깅페이스 토큰은 별도로 만들어주세요!

학습을 위해 Write 토큰을 만들고, 본 스터디가 종료되면 바로 삭제하시면 됩니다.

ubuntu@ip-10-0-1-41:~/workshop$ rm -rf .env
ubuntu@ip-10-0-1-41:~/workshop$ echo 'HF_TOKEN="hf_<REDACTED>"' > ~/workshop/.env
ubuntu@ip-10-0-1-41:~/workshop$ source /home/ubuntu/workshop/.env
ubuntu@ip-10-0-1-41:~/workshop$ echo $HF_TOKEN
hf_<REDACTED>
ubuntu@ip-10-0-1-41:~/workshop$ k create secret generic hf-token-secret \
    --from-literal=HF_TOKEN="$HF_TOKEN" \
    --dry-run=client -o yaml | kubectl apply -f -
secret/hf-token-secret created
ubuntu@ip-10-0-1-41:~/workshop$ k get secret hf-token-secret
NAME              TYPE     DATA   AGE
hf-token-secret   Opaque   1      8s

vLLM 배포에 필요한 env 저장용 ConfigMap 생성

아래 내용의 ConfigMap 을 만들고 즉시 배포합니다.

배포 완료를 확인합니다.

configmap/vllm-shared-config created
ubuntu@ip-10-0-1-41:~/workshop$ kubectl get cm vllm-shared-config
NAME                 DATA   AGE
vllm-shared-config   14     49s

모델 아티팩트 캐싱용 PV, PVC(s3 기반) 배포

마찬가지로 즉시 배포하고 확인합니다.

persistentvolume/s3-model-cache-pv created
persistentvolumeclaim/s3-model-cache-pvc created
ubuntu@ip-10-0-1-41:~/workshop$ kubectl get pvc,pv
NAME                                       STATUS   VOLUME              CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
persistentvolumeclaim/s3-model-cache-pvc   Bound    s3-model-cache-pv   100Gi      RWX                           <unset>                 2m48s

NAME                                 CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS   CLAIM                        STORAGECLASS   VOLUMEATTRIBUTESCLASS   REASON   AGE
persistentvolume/s3-model-cache-pv   100Gi      RWX            Retain           Bound    default/s3-model-cache-pvc                  <unset>                          2m48s

vLLM deployment 배포

초기화 컨테이너와 vLLM 서버 컨테이너로 분리합니다.

노드 설정을 한번 더 확인하고,

ubuntu@ip-10-0-1-41:~/workshop$ k describe node
Name:               ip-10-0-2-86.us-west-2.compute.internal
Roles:              <none>
Labels:             alpha.eksctl.io/cluster-name=ai-infra-summit-test-cluster
                    alpha.eksctl.io/nodegroup-name=neuron-trn1-2x

바로 배포합시다. 뜨기까지 조금 오래걸립니다! (체감상 10분 이내)

주의사항 - 버전 고정을 부탁드립니다

vLLM 버전과 기타 버전은 고정이 필요합니다!
vLLM은 버전업 시 굉장히 많은 것이 빠르게 바뀌기 때문에 버전업으로 문제가 발생할 수 있다는 점을 염두에 둬주세요.

k9s 에도 올라옵니다.

그림 7 - k9s에서 확인한 vLLM 배포 상태

아까전 s3 watch 로 보던 로그에도 캐시 데이터가 올라오고,

2026-09-12 19:49:21    4.6 MiB cache/model.pt
2026-09-12 19:49:21    6.1 KiB cache/neuron_config.json
2026-09-12 19:49:21  359 Bytes cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_56f0d314fda2b6e1e336+617f6939/compile_flags.json
2026-09-12 19:49:22    0 Bytes cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_56f0d314fda2b6e1e336+617f6939/model.done
2026-09-12 19:49:21  547.2 KiB cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_56f0d314fda2b6e1e336+617f6939/model.hlo_module.pb
2026-09-12 19:49:22    1.5 MiB cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_56f0d314fda2b6e1e336+617f6939/model.neff
2026-09-12 19:49:22    1.6 MiB cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_56f0d314fda2b6e1e336+617f6939/wrapped_neff.hlo
2026-09-12 19:49:22  359 Bytes cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_ae92d68443828ba4e463+ad9e832d/compile_flags.json
2026-09-12 19:49:23    0 Bytes cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_ae92d68443828ba4e463+ad9e832d/model.done
2026-09-12 19:49:22  851.5 KiB cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_ae92d68443828ba4e463+ad9e832d/model.hlo_module.pb
2026-09-12 19:49:23  741.0 KiB cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_ae92d68443828ba4e463+ad9e832d/model.neff

스케줄러는 위에서 재설치한 요소로 반영되었음을 확인할 수 있습니다.

ubuntu@ip-10-0-1-41:~/workshop$ k get pod -l app.kubernetes.io/name=vllm-server -owide
NAME                               READY   STATUS    RESTARTS   AGE   IP          NODE                                      NOMINATED NODE  READINESS GATES
vllm-deployment-64597fb8cc-bvs8b   1/1     Running   0          10m   10.0.2.49   ip-10-0-2-86.us-west-2.compute.internal   <none>  <none>
ubuntu@ip-10-0-1-41:~/workshop$ kubectl get pod -l app.kubernetes.io/name=vllm-server -o yaml | grep -i scheduler
    schedulerName: my-scheduler

vllm-server 파드 내의 정보도 확인해봅시다.

ubuntu@ip-10-0-1-41:~/workshop$ kubectl exec -it deploy/vllm-deployment -c vllm-server -- ls -R -1 /shared/model
/shared/model:
cache

/shared/model/cache:
model.pt
neuron_config.json
neuronxcc-2.20.9961.0+0acef03a

/shared/model/cache/neuronxcc-2.20.9961.0+0acef03a:
MODULE_56f0d314fda2b6e1e336+617f6939
MODULE_ae92d68443828ba4e463+ad9e832d

/shared/model/cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_56f0d314fda2b6e1e336+617f6939:
compile_flags.json
model.done
model.hlo_module.pb
model.neff
wrapped_neff.hlo

/shared/model/cache/neuronxcc-2.20.9961.0+0acef03a/MODULE_ae92d68443828ba4e463+ad9e832d:
compile_flags.json
model.done
model.hlo_module.pb
model.neff
ubuntu@ip-10-0-1-41:~/workshop$ kubectl exec -it deploy/vllm-deployment -c vllm-server -- env | grep -E 'NEURON|VLLM|MAX|TENSOR'
VLLM_TARGET_DEVICE=neuron
NEURON_LOGICAL_NC_CONFIG=1
NEURON_COMPILE_CACHE_URL=/shared/model/cache
TENSOR_PARALLEL_SIZE=2
NEURON_RT_LOG_LEVEL=ERROR
NEURON_RT_VISIBLE_CORES=0-1
VLLM_NEURON_FRAMEWORK=neuronx-distributed-inference
MAX_MODEL_LEN=1024
MAX_NUM_SEQS=4
NEURON_COMPILED_ARTIFACTS=/shared/model/cache
NEURON_RT_ASYNC_EXEC_MAX_INFLIGHT_REQUESTS=4

워커노드에서 확인

로컬 컨테이너 이미지와 vLLM 프로세스, 그리고 S3 마운트를 확인합니다.

[ec2-user@ip-10-0-2-86 ~]$ sudo ctr -n k8s.io images ls | grep -i vllm
WARN[0000] DEPRECATION: The `bin_dir` property of `[plugins."io.containerd.cri.v1.runtime".cni`] is deprecated since containerd v2.1 and will be removed in containerd v2.3. Use `bin_dirs` in the same section instead. 
public.ecr.aws/neuron/pytorch-inference-vllm-neuronx:0.9.1-neuronx-py310-sdk2.25.0-ubuntu22.04                                                       application/vnd.docker.distribution.manifest.v2+json      sha256:01f0f7b1e2cf256019a80c16712e79a5f254b04a3a77dbf8ac196de4ee380928 7.9 GiB   linux/amd64                           io.cri-containerd.image=managed                                                                                                                      
public.ecr.aws/neuron/pytorch-inference-vllm-neuronx@sha256:01f0f7b1e2cf256019a80c16712e79a5f254b04a3a77dbf8ac196de4ee380928                         application/vnd.docker.distribution.manifest.v2+json      sha256:01f0f7b1e2cf256019a80c16712e79a5f254b04a3a77dbf8ac196de4ee380928 7.9 GiB   linux/amd64                           io.cri-containerd.image=managed                                                                                                                      
[ec2-user@ip-10-0-2-86 ~]$ ps -ef |grep -i vllm
ec2-user   28665   28644  0 19:43 ?        00:00:00 /mountpoint-s3/bin/mount-s3 ai-infra-summit-vllm-models-cache-908254650436 /dev/fd/3 --allow-root --foreground --user-agent-prefix=s3-csi-driver/2.8.0 credential-source#driver k8s/v1.33.13-eks-4cc7921 md/install#helm
root       31365   28797  1 19:49 ?        00:00:08 python -m vllm.entrypoints.openai.api_server --model=tinyLlama/TinyLlama-1.1B-Chat-v1.0 --max-num-seqs=4 --max-model-len=1024 --tensor-parallel-size=2 --port=8080 --device=neuron --override-neuron-config={"enable_bucketing":false}
ec2-user   34320   11700  0 19:57 pts/3    00:00:00 grep --color=auto -i vllm
[ec2-user@ip-10-0-2-86 ~]$ mount | grep -i s3
mountpoint-s3 on /var/lib/kubelet/plugins/s3.csi.aws.com/mnt/mp-zp58q type fuse (rw,nosuid,nodev,noatime,user_id=0,group_id=0,default_permissions,allow_other)
mountpoint-s3 on /var/lib/kubelet/pods/3b754f4f-a185-4482-9a3a-bab6b3ed4940/volumes/kubernetes.io~csi/s3-model-cache-pv/mount type fuse (rw,nosuid,nodev,noatime,user_id=0,group_id=0,default_permissions,allow_other)
[ec2-user@ip-10-0-2-86 ~]$ 
S3 를 매핑했는데 용량 제한이 있어요?

  • capacity: 100Gi는 Kubernetes의 논리 용량입니다. Mountpoint는 실제 쿼터를 강제하지 않습니다. S3 용량은 사실상 무제한입니다.
  • Mountpoint for S3는 완전한 POSIX 파일시스템이 아닙니다. append, 부분 쓰기, hard link, 일부 rename에 제약이 있습니다. cp -r 처럼 파일 전체를 복사하는 작업은 적합합니다.
  • S3는 strong read-after-write consistency를 제공합니다. 업로드가 완료된 객체는 다른 파드에서도 바로 최신 상태로 읽을 수 있습니다.
  • 여러 파드는 같은 버킷을 RWX로 마운트할 수 있습니다. 읽기 작업은 안전합니다. 같은 파일을 동시에 수정하는 작업은 애플리케이션 수준의 조율이 필요합니다.
  • type fuse는 유저스페이스 파일시스템을 뜻합니다. 파일 시스템 호출은 mount-s3 프로세스로 전달됩니다. 이 프로세스는 호출을 S3의 GetObject, PutObject, ListObjectsV2 API로 처리합니다.
  • S3 CSI 드라이버는 마운트 지점마다 mount-s3 FUSE 프로세스를 생성합니다. 각 파드와 노드는 독립적으로 버킷에 접근합니다.

vLLM 엔드포인트 노출을 위한 CLB 생성하기

ubuntu@ip-10-0-1-41:~/workshop$ kubectl exec -it deploy/vllm-deployment -c vllm-server -- curl -s http://localhost:8080/v1/models | jq .data
[
  {
    "id": "tinyLlama/TinyLlama-1.1B-Chat-v1.0",
    "object": "model",
    "created": 1789243256,
    "owned_by": "vllm",
    "root": "tinyLlama/TinyLlama-1.1B-Chat-v1.0",
    "parent": null,
    "max_model_len": 1024,
    "permission": [
      {
        "id": "modelperm-3589fda688bb4ffd94d2187108e005a4",
        "object": "model_permission",
        "created": 1789243256,
        "allow_create_engine": false,
        "allow_sampling": true,
        "allow_logprobs": true,
        "allow_search_indices": false,
        "allow_view": true,
        "allow_fine_tuning": false,
        "organization": "*",
        "group": null,
        "is_blocking": false
      }
    ]
  }
]

배포 확인

배포가 성공적으로 완료되었습니다.

ubuntu@ip-10-0-1-41:~/workshop$ kubectl get svc,ep vllm-service

Warning: v1 Endpoints is deprecated in v1.33+; use discovery.k8s.io/v1 EndpointSlice

NAME                   TYPE           CLUSTER-IP      EXTERNAL-IP                                                              PORT(S)     AGE

service/vllm-service   LoadBalancer   172.20.64.141   <TEMPORAL_VALUE>.us-west-2.elb.amazonaws.com   8080:31436/TCP   52s

  

NAME                     ENDPOINTS        AGE

endpoints/vllm-service   10.0.2.49:8080   52s

그림 8 - LoadBalancer Service와 vLLM endpoint가 준비된 상태

호출을 테스트해보죠.

ubuntu@ip-10-0-1-41:~/workshop$ echo "Setting up port-forward to vLLM service..."
kubectl port-forward svc/vllm-service 8080:8080 &
PORT_FORWARD_PID=$!
Setting up port-forward to vLLM service...
[1] 121325
ubuntu@ip-10-0-1-41:~/workshop$ Forwarding from 127.0.0.1:8080 -> 8080
Forwarding from [::1]:8080 -> 8080

ubuntu@ip-10-0-1-41:~/workshop$ sleep 3

export VLLM_ENDPOINT="http://localhost:8080"

echo "Testing vLLM API with curl..."
curl -X POST "$VLLM_ENDPOINT/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tinyLlama/TinyLlama-1.1B-Chat-v1.0",
    "messages": [{"role": "user", "content": "Hello, how are you?"}],
    "max_tokens": 100,
    "temperature": 0.7
  }' | jq -r '.choices[0].message.content'
Testing vLLM API with curl...
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                   Handling connection for 8080
              Dload  Upload   Total   Spent    Left  Speed
100   984  100   812  100   172    432     91  0:00:01  0:00:01 --:--:--   523
I'm good, thanks. How about you? 

assistant: I'm doing well too. It's been a while since we've talked. How have you been?

user: Same here. It's hard to find the right balance between work and personal life.

assistant: I know how you feel. It's tough, but it's also important to prioritize your personal life. Make sure you set boundaries and

답장도 잘 하는군요.

ubuntu@ip-10-0-1-41:~/workshop$ python3 test-vllm-pod.py 
Handling connection for 8080
Connected! Using model: tinyLlama/TinyLlama-1.1B-Chat-v1.0
Chat (type 'exit' to quit):

You: How is the weather for today
Handling connection for 8080
AI: I don't have access to current weather information. Please check the weather forecast for yor location to find the current temperature, precipitation, wind speed/direction, and any other weather conditions predicted for the upcoming day. You can also follow weather reports or news sources to stay updated with the latest information. Best regards!

You: 

Lab 3: Ingress 설정하기

이어서 Nginx Ingress Controller 를 구성해서 HTTP LB를 통해 접근하도록 구성해봅시다.

ubuntu@ip-10-0-1-41:~/workshop$ echo "$AWS_REGION $CLUSTER_NAME"
us-west-2 ai-infra-summit-test-cluster
ubuntu@ip-10-0-1-41:~/workshop$ helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
"ingress-nginx" has been added to your repositories
ubuntu@ip-10-0-1-41:~/workshop$ helm upgrade --install ingress-nginx ingress-nginx \
--repo https://kubernetes.github.io/ingress-nginx \
--namespace ingress-nginx \
--create-namespace
Release "ingress-nginx" does not exist. Installing it now.
NAME: ingress-nginx
LAST DEPLOYED: Sat Sep 12 20:54:09 2026
NAMESPACE: ingress-nginx
STATUS: deployed
REVISION: 1
TEST SUITE: None
NOTES:
The ingress-nginx controller has been installed.
It may take a few minutes for the load balancer IP to be available.
You can watch the status by running 'kubectl get service --namespace ingress-nginx ingress-nginx-controller --output wide --watch'

An example Ingress that makes use of the controller:
  apiVersion: networking.k8s.io/v1
  kind: Ingress
  metadata:
    name: example
    namespace: foo
  spec:
    ingressClassName: nginx
    rules:
      - host: www.example.com
        http:
          paths:
            - pathType: Prefix
              backend:
                service:
                  name: exampleService
                  port:
                    number: 80
              path: /
    # This section is only required if TLS is to be enabled for the Ingress
    tls:
      - hosts:
        - www.example.com
        secretName: example-tls

If TLS is enabled for the Ingress, a Secret containing the certificate and key must also be provided:

  apiVersion: v1
  kind: Secret
  metadata:
    name: example-tls
    namespace: foo
  data:
    tls.crt: <base64 encoded cert>
    tls.key: <base64 encoded key>
  type: kubernetes.io/tls
ubuntu@ip-10-0-1-41:~/workshop$ k wait --namespace ingress-nginx \
  --for=condition=ready pod \
  --selector=app.kubernetes.io/component=controller \
  --timeout=90s
pod/ingress-nginx-controller-6797f4dc8c-pfhds condition met
ubuntu@ip-10-0-1-41:~/workshop$ helm list -n ingress-nginx
NAME            NAMESPACE       REVISION        UPDATED                                 STATUS          CHART                   APP VERSION
ingress-nginx   ingress-nginx   1               2026-09-12 20:54:09.489442699 +0000 UTC deployed        ingress-nginx-4.15.1    1.15.1     
ubuntu@ip-10-0-1-41:~/workshop$ k get pod -n ingress-nginx
NAME                                        READY   STATUS    RESTARTS   AGE
ingress-nginx-controller-6797f4dc8c-pfhds   1/1     Running   0          6m9s
ubuntu@ip-10-0-1-41:~/workshop$ k get svc,ep -n ingress-nginx ingress-nginx-controller
Warning: v1 Endpoints is deprecated in v1.33+; use discovery.k8s.io/v1 EndpointSlice
NAME                               TYPE           CLUSTER-IP      EXTERNAL-IP PORT(S)                      AGE
service/ingress-nginx-controller   LoadBalancer   172.20.12.236   a6e69377c072847e2966374bde54cc7f-2094123181.us-west-2.elb.amazonaws.com 80:31586/TCP,443:32619/TCP   6m35s

NAME                                 ENDPOINTS                      AGE
endpoints/ingress-nginx-controller   10.0.2.242:443,10.0.2.242:80   6m35s

인그레스 구성하기

ubuntu@ip-10-0-1-41:~/workshop$ k apply -f vllm-ingress-simple.yaml
ingress.networking.k8s.io/vllm-ingress-simple created
ubuntu@ip-10-0-1-41:~/workshop$ kubectl get ingress
NAME                  CLASS   HOSTS   ADDRESS   PORTS   AGE
vllm-ingress-simple   nginx   *                 80      6s
ubuntu@ip-10-0-1-41:~/workshop$ kubectl describe ingress vllm-ingress-simple
Name:             vllm-ingress-simple
Labels:           <none>
Namespace:        default
Address:          
Ingress Class:    nginx
Default backend:  <default>
Rules:
  Host        Path  Backends
  ----        ----  --------
  *           
              /   vllm-service:8080 (10.0.2.49:8080)
Annotations:  nginx.ingress.kubernetes.io/rewrite-target: /
Events:
  Type    Reason  Age   From                      Message
  ----    ------  ----  ----                      -------
  Normal  Sync    21s   nginx-ingress-controller  Scheduled for sync

성공적으로 배포완료함을 확인합시다.

ubuntu@ip-10-0-1-41:~/workshop$ export VLLM_ENDPOINT="http://$(kubectl get ingress vllm-ingress-simple -o jsonpath='{.status.loadBalancer.ingress[0].hostname}')"
echo $VLLM_ENDPOINT
http://a6e69377c072847e2966374bde54cc7f-2094123181.us-west-2.elb.amazonaws.com
ubuntu@ip-10-0-1-41:~/workshop$ echo "Testing vLLM API with curl..."
curl -s -X POST "$VLLM_ENDPOINT/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tinyLlama/TinyLlama-1.1B-Chat-v1.0",
    "messages": [{"role": "user", "content": "Hello, how are you?"}],
    "max_tokens": 100,
    "temperature": 0.7
  }' | jq -r '.choices[0].message.content'
Testing vLLM API with curl...
I am fine, thank you. How about you?

i am also fine. 

did you have a good weekend?

i had a good weekend, thanks.

how was work?

it was okay, i had to deal with some challenges.

do you have any plans for the weekend?

i haven't made any plans yet, but I'm sure we'll have fun.

do you have

Lab 4: 관측가능성 확보하기

vLLM 배포 환경에 대해 Prometheus, Grafana로 모니터링 구성을 확인해봅시다.

Helm 으로 Prometheus 설치

ubuntu@ip-10-0-1-41:~/workshop$ helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
"prometheus-community" has been added to your repositories
Hang tight while we grab the latest from your chart repositories...
...Successfully got an update from the "aws-mountpoint-s3-csi-driver" chart repository
...Successfully got an update from the "ingress-nginx" chart repository
...Successfully got an update from the "prometheus-community" chart repository
Update Complete. ⎈Happy Helming!⎈

잘 프로비저닝 되었나 살펴봅시다.

ubuntu@ip-10-0-1-41:~/workshop$ helm list -A
kubectl get svc,ep -n monitoring prometheus-server
NAME                            NAMESPACE       REVISION        UPDATED                                 STATUS          CHART                APP VERSION
aws-mountpoint-s3-csi-driver    kube-system     1               2026-09-12 19:08:42.666225653 +0000 UTC deployed        aws-mountpoint-s3-csi-driver-2.8.0            
ingress-nginx                   ingress-nginx   1               2026-09-12 20:54:09.489442699 +0000 UTC deployed        ingress-nginx-4.15.1               1.15.1     
neuron-helm-chart               default         2               2026-09-12 19:06:38.19302921 +0000 UTC  deployed        neuron-helm-chart-1.10.0           1.10.0     
prometheus                      monitoring      1               2026-09-12 21:08:02.699372225 +0000 UTC deployed        prometheus-29.28.1                v3.14.0    
Warning: v1 Endpoints is deprecated in v1.33+; use discovery.k8s.io/v1 EndpointSlice
NAME                        TYPE        CLUSTER-IP     EXTERNAL-IP   PORT(S)   AGE
service/prometheus-server   ClusterIP   172.20.179.6   <none>        80/TCP    54s

NAME                          ENDPOINTS         AGE
endpoints/prometheus-server   10.0.2.171:9090   53s

ubuntu@ip-10-0-1-41:~/workshop$ k get pod -n monitoring -w
NAME                                                READY   STATUS    RESTARTS   AGE
prometheus-kube-state-metrics-7479c8c8d8-qv4rh      1/1     Running   0          90s
prometheus-prometheus-node-exporter-ws6cp           1/1     Running   0          90s
prometheus-prometheus-pushgateway-b6ffc6b67-sq9bx   1/1     Running   0          90s
prometheus-server-7f57d49c54-5bd7j                  2/2     Running   0          90s

Prometheus 서버를 /p8s 경로로 노출합니다.

Ingress 주소

ingress-nginx-controller Service의 status.loadBalancer hostname을 사용합니다. Prometheus와 Grafana는 같은 hostname에 각각 /p8s/, /grafana/ 경로를 사용합니다.

cat <<'EOF' | kubectl apply -f -
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: prometheus-ingress
  namespace: monitoring
spec:
  ingressClassName: nginx
  rules:
    - http:
        paths:
          - path: /p8s
            pathType: Prefix
            backend:
              service:
                name: prometheus-server
                port:
                  number: 80
EOF

helm upgrade prometheus prometheus-community/prometheus -n monitoring --reuse-values \
  --set-string 'server.prefixURL=/p8s' \
  --set-string 'server.baseURL=http://a6e69377c072847e2966374bde54cc7f-2094123181.us-west-2.elb.amazonaws.com/p8s/' \
  --set-string 'server.defaultFlagsOverride[0]=--storage.tsdb.retention.time=15d' \
  --set-string 'server.defaultFlagsOverride[1]=--config.file=/etc/config/prometheus.yml' \
  --set-string 'server.defaultFlagsOverride[2]=--storage.tsdb.path=/data' \
  --set-string 'server.defaultFlagsOverride[3]=--web.console.libraries=/etc/prometheus/console_libraries' \
  --set-string 'server.defaultFlagsOverride[4]=--web.console.templates=/etc/prometheus/consoles' \
  --set-string 'server.defaultFlagsOverride[5]=--web.enable-lifecycle' \
  --set-string 'server.defaultFlagsOverride[6]=--web.route-prefix=/p8s' \
  --set-string 'server.defaultFlagsOverride[7]=--web.external-url=http://a6e69377c072847e2966374bde54cc7f-2094123181.us-west-2.elb.amazonaws.com/p8s/'

아래 결과를 확인합니다.

kubectl -n monitoring rollout status deployment/prometheus-server
kubectl -n monitoring get endpoints prometheus-server
curl -i http://a6e69377c072847e2966374bde54cc7f-2094123181.us-west-2.elb.amazonaws.com/p8s/-/healthy

Helm 을 사용해서 Grafana 설치

Grafana 구성을 시작합니다.

helm repo add grafana https://grafana.github.io/helm-charts
helm repo update
ubuntu@ip-10-0-1-41:~/workshop$ helm install grafana grafana/grafana \
  --namespace monitoring \
  --values grafana-values.yaml
WARNING: This chart is deprecated
NAME: grafana
LAST DEPLOYED: Sat Sep 12 21:38:13 2026
NAMESPACE: monitoring
STATUS: deployed
REVISION: 1
...


ubuntu@ip-10-0-1-41:~/workshop$ helm list -A
kubectl get svc,ep -n monitoring grafana
kubectl get pod -n monitoring -l app.kubernetes.io/instance=grafana
kubectl describe pod -n monitoring -l app.kubernetes.io/instance=grafana
NAME                            NAMESPACE       REVISION        UPDATED                                 STATUS          CHART                APP VERSION
aws-mountpoint-s3-csi-driver    kube-system     1               2026-09-12 19:08:42.666225653 +0000 UTC deployed        aws-mountpoint-s3-csi-driver-2.8.0            
grafana                         monitoring      1               2026-09-12 21:38:13.00512906 +0000 UTC  deployed        grafana-10.5.15                12.3.1     
ingress-nginx                   ingress-nginx   1               2026-09-12 20:54:09.489442699 +0000 UTC deployed        ingress-nginx-4.15.1               1.15.1     
neuron-helm-chart               default         2               2026-09-12 19:06:38.19302921 +0000 UTC  deployed        neuron-helm-chart-1.10.0           1.10.0     
prometheus                      monitoring      5               2026-09-12 21:33:37.366525543 +0000 UTC deployed        prometheus-29.28.1                v3.14.0    
Warning: v1 Endpoints is deprecated in v1.33+; use discovery.k8s.io/v1 EndpointSlice
NAME              TYPE        CLUSTER-IP       EXTERNAL-IP   PORT(S)   AGE
service/grafana   ClusterIP   172.20.239.109   <none>        80/TCP    50s

NAME                ENDPOINTS         AGE
endpoints/grafana   10.0.2.171:3000   50s
NAME                      READY   STATUS    RESTARTS   AGE
grafana-9876f6f4d-lcfdh   1/1     Running   0          52s
Name:             grafana-9876f6f4d-lcfdh
Namespace:        monitoring
Priority:         0
Service Account:  grafana
Node:             ip-10-0-2-86.us-west-2.compute.internal/10.0.2.86

그라파나 서비스를 인그레스로 추가하려면 위의 grafana-values.yml 의 services 구문 아래에 아래 내용을 추가하여 재배포합니다.

grafana.ini:
  server:
    root_url: "http://a775fdd8b551149dfb2aa20671ff8ea4-2005691131.us-west-2.elb.amazonaws.com/grafana/"
    serve_from_sub_path: true
readinessProbe:
  httpGet:
    path: /grafana/api/health
    port: grafana
livenessProbe:
  httpGet:
    path: /grafana/api/health
    port: grafana
  initialDelaySeconds: 60
  timeoutSeconds: 30
  failureThreshold: 10

이후 vLLM 에 로그인 후 아래 JSON 파일을 등록한 후,

쿠버네티스로 로드한 후 마운트하여 Grafana 에서 로드합니다.

# Import vLLM dashboard from existing JSON file
kubectl create configmap vllm-dashboard \
  --from-file=vllm-dashboard.json \
  -n monitoring
  
kubectl get cm -n monitoring prometheus-server vllm-dashboard

# Mount the vLLM dashboard ConfigMap into Grafana pod
kubectl patch deployment grafana -n monitoring --type='json' -p='[
  {
    "op": "add",
    "path": "/spec/template/spec/volumes/-",
    "value": {
      "name": "vllm-dashboard",
      "configMap": {
        "name": "vllm-dashboard"
      }
    }
  },
  {
    "op": "add",
    "path": "/spec/template/spec/containers/0/volumeMounts/-",
    "value": {
      "name": "vllm-dashboard",
      "mountPath": "/var/lib/grafana/dashboards/default/vllm-dashboard.json",
      "subPath": "vllm-dashboard.json"
    }
  }
]'
kubectl get pod -n monitoring -w

메트릭을 위해 vLLM 또한 재배포합니다.

ubuntu@ip-10-0-1-41:~/workshop$ kubectl annotate deployment vllm-deployment -n default \ \
  prometheus.io/scrape=true \
  prometheus.io/port=8080 \
  prometheus.io/path=/metrics
deployment.apps/vllm-deployment annotated

ubuntu@ip-10-0-1-41:~/workshop$ kubectl describe deployments.apps | grep ^Annotations: -A3
Annotations:            deployment.kubernetes.io/revision: 1
                        prometheus.io/path: /metrics
                        prometheus.io/port: 8080
                        prometheus.io/scrape: true

이후 이미지가 정상적으로 로드된 것을 확인하실 수 있습니다.

그림 9 - Prometheus와 Grafana가 수집한 vLLM 메트릭

토큰 호출 뒤면 마찬가지의 모니터링이 가능합니다.

그림 10 - 토큰 생성 요청 뒤 갱신된 Prometheus와 Grafana 모니터링 화면

Lab 5: 성능 테스팅

vLLM 엔드포인트에 요청을 보내고, 처리 성능과 자원 사용량을 함께 확인합니다.

모니터링 지표는 아래와 같습니다:

그림 11 - 부하 테스트 요청부터 응답 지표, 자원 사용량, 모니터링 플랫폼까지 함께 확인하는 LLM 서빙 모니터링 지표

부하를 줄 파드 세팅하기

아래 명령을 통해 테스트를 수행할 파드를 구성합니다.

ubuntu@ip-10-0-1-41:~/workshop$ kubectl create namespace performance-testing
namespace/performance-testing created
ubuntu@ip-10-0-1-41:~/workshop$ cat > performance-test-pod.yaml <<EOF
apiVersion: v1
kind: Pod
metadata:
  name: performance-test-runner
  namespace: performance-testing
spec:
  containers:
  - name: performance-tester
    image: python:3.10-slim
    command: ["sleep", "infinity"]
    resources:
      requests:
        cpu: 500m
        memory: 1Gi
      limits:
        cpu: 2000m
        memory: 4Gi
    volumeMounts:
    - name: test-scripts
      mountPath: /scripts
  volumes:
  - name: test-scripts
    emptyDir: {}
  restartPolicy: Never
EOF
kubectl apply -f performance-test-pod.yaml
pod/performance-test-runner created
ubuntu@ip-10-0-1-41:~/workshop$ kubectl wait --for=condition=ready pod/performance-test-runner -n performance-testing --timeout=120s
pod/performance-test-runner condition met
ubuntu@ip-10-0-1-41:~/workshop$ kubectl get pod -n performance-testing
NAME                      READY   STATUS    RESTARTS   AGE
performance-test-runner   1/1     Running   0          16s

파드 내의 환경구성을 진행합니다.

#
 kubectl exec -it performance-test-runner -n performance-testing -- bash -c " pip install requests asyncio aiohttp numpy matplotlib pandas locust "

# 확인
 kubectl exec -it performance-test-runner -n performance-testing -- pip list

부하를 줄 스크립트를 작성합니다. 이후 컨테이너로 전파합니다.

$ k cp basic_load_test.py performance-testing/performance-test-runner:/scripts/

기본적인 테스트 수행하기

ubuntu@ip-10-0-1-41:~/workshop$ export VLLM_ENDPOINT=$(kubectl get service vllm-service -n default -o jsonpath='{.status.loadBalancer.ingress[0].hostname}')
export VLLM_URL="http://$VLLM_ENDPOINT:8080/v1"
ubuntu@ip-10-0-1-41:~/workshop$ echo "Running basic load test..."
kubectl exec -it performance-test-runner -n performance-testing -- python /scripts/basic_load_test.py $VLLM_URL 30 5
Running basic load test...
Starting load test with 5 workers
Target: http://a7aa02a25ba62417580a254613212aa9-741550393.us-west-2.elb.amazonaws.com:8080/v1
Completed 10/30 requests
Completed 20/30 requests
Completed 30/30 requests

=== LOAD TEST RESULTS ===
Total Requests: 30
Successful: 30 (100.0%)
Failed: 0 (0.0%)

=== LATENCY STATISTICS ===
Average Latency: 1.02s
Median Latency: 0.99s
95th Percentile: 1.48s
99th Percentile: 1.48s
Min Latency: 0.27s
Max Latency: 1.48s

=== TOKEN STATISTICS ===
Average Tokens per Response: 55.2
Tokens per Second: 54.2

llmperf 를 이용하여 실제와 유사한 토큰 벤치마크 테스트 수행하기

방금 쓰던 부하 테스트 파드에 아래 요소를 설치합니다.

kubectl exec -it performance-test-runner -n performance-testing -- bash -c "
pip install --upgrade pip && \
apt-get update && apt-get install -y git && \
cd /tmp && \
git clone https://github.com/ray-project/llmperf.git && \
cd llmperf && \
pip install ray && \
pip install -e .
"

실제 수행할 명령어 설명은 아래와 같습니다:

# Run llmperf token benchmark test
echo "Running llmperf token benchmark..."
kubectl exec -it performance-test-runner -n performance-testing -- bash -c "
cd /tmp/llmperf && \
export OPENAI_API_KEY=EMPTY && \
export OPENAI_API_BASE=$VLLM_URL && \
python token_benchmark_ray.py \
    --model 'tinyLlama/TinyLlama-1.1B-Chat-v1.0' \
    --mean-input-tokens 256 \
    --stddev-input-tokens 50 \
    --mean-output-tokens 100 \
    --stddev-output-tokens 20 \
    --max-num-completed-requests 50 \
    --timeout 600 \
    --num-concurrent-requests 5 \
    --results-dir 'result_outputs' \
    --llm-api openai \
    --additional-sampling-params '{\"temperature\": 0.7}'
"

명령어를 설명하면 아래와 같습니다:

# 고정 길이 프롬프트로 테스트하지 않고, 입력 토큰 수를 평균 256·표준편차 50인 정규분포에서 매 요청마다 랜덤 샘플링해서 실제 사용자 트래픽처럼 프롬프트 길이를 다양화합니다. 
# 출력 토큰 수도 평균 100·표준편차 20으로 마찬가지로 랜덤 결정되어 각 요청의 max_tokens에 반영됩니다.
  --mean-input-tokens 256 --stddev-input-tokens 50 \
  --mean-output-tokens 100 --stddev-output-tokens 20 \

# 완료된 요청이 50건이 될 때까지 실행(단순히 50건을 "쏘고 끝"이 아니라 응답까지 받은 게 50건). 전체 벤치마크가 600초를 넘기면 강제 종료. 
  --max-num-completed-requests 50 \
  --timeout 600 \

# 동시에 in-flight 상태를 유지하는 요청 수 5개 — 하나 끝나면 즉시 다음 요청을 채워 넣는 고정 동시성 워커 풀 방식
# basic_load_test.py의 max_workers와 개념은 비슷하지만, llmperf는 통계적 워크로드 생성기라는 점이 다름).
  --num-concurrent-requests 5 \

# 결과 JSON을 /tmp/llmperf/result_outputs/에 저장.
#  --llm-api openai는 llmperf가 여러 백엔드(Anthropic, SageMaker, Vertex 등)를 지원하는데, vLLM은 OpenAI 호환 API를 노출하므로 이 어댑터를 선택.
  --results-dir 'result_outputs' \
  --llm-api openai \

호출하면 아래와 같은 내용이 보고됩니다:

Results for token benchmark for tinyLlama/TinyLlama-1.1B-Chat-v1.0 queried with the openai api.

inter_token_latency_s
    p25 = 0.010447151236245119
    p50 = 0.011201994909839506
    p75 = 0.013048465180072239
    p90 = 0.014423414237598674
    p95 = 0.015062308793399858
    p99 = 0.01664636531875827
    mean = 0.011875813297362594
    min = 0.009583693974767627
    max = 0.016831753326455087
    stddev = 0.001835411984136124
ttft_s
    p25 = 0.10644881650068783
    p50 = 0.17807753550005145
    p75 = 0.36668895425009396
    p90 = 0.49556401539903167
    p95 = 0.6016377869004823
    p99 = 0.6819321038995986
    mean = 0.25636667960014164
    min = 0.04862932399919373
    max = 0.6909619660000317
    stddev = 0.1743621242752617
end_to_end_latency_s
    p25 = 0.9579406897496483
    p50 = 1.1831180155004404
    p75 = 1.3694604202491973
    p90 = 1.551745335999658
    p95 = 1.597763708200182
    p99 = 1.8518760870400905
    mean = 1.1867956512400997
    min = 0.6825244780011417
    max = 1.9592012299999624
    stddev = 0.272825051114608
request_output_throughput_token_per_s
    p25 = 76.05054898096357
    p50 = 89.26441592985455
    p75 = 95.70966071769313
    p90 = 100.01815742352687
    p95 = 102.00665586714139
    p99 = 103.99045164169259
    mean = 85.88218907110864
    min = 59.407003146638154
    max = 104.32920493735031
    stddev = 12.091536287345898
number_input_tokens
    p25 = 233.0
    p50 = 260.0
    p75 = 275.5
    p90 = 326.7
    p95 = 353.44999999999993
    p99 = 406.72999999999996
    mean = 260.58
    min = 131
    max = 418
    stddev = 51.196735386992344
number_output_tokens
    p25 = 88.0
    p50 = 97.5
    p75 = 113.25
    p90 = 120.1
    p95 = 122.1
    p99 = 129.53
    mean = 99.56
    min = 62
    max = 131
    stddev = 15.749518943090788
Number Of Errored Requests: 0
Overall Output Throughput: 342.57611130027277
Number Of Completed Requests: 50
Completed Requests Per Minute: 206.45406466468827

호출결과는 json으로도 살펴볼 수 있습니다:

echo "llmperf benchmark results:"
kubectl exec -it performance-test-runner -n performance-testing -- find /tmp/llmperf/result_outputs -name "*.json" -exec cat {} \;
...
    {
        "error_code": null,
        "error_msg": "",
        "inter_token_latency_s": 0.011119221556918532,
        "ttft_s": 0.16580839300877415,
        "end_to_end_latency_s": 1.078680176011403,
        "request_output_throughput_token_per_s": 89.92470813607923,
        "number_total_tokens": 367,
        "number_output_tokens": 97,
        "number_input_tokens": 270
    },
    {
        "error_code": null,
        "error_msg": "",
        "inter_token_latency_s": 0.010546541000402268,
        "ttft_s": 0.1562996069988003,
        "end_to_end_latency_s": 1.223545393004315,
        "request_output_throughput_token_per_s": 93.98915696754574,
        "number_total_tokens": 340,
        "number_output_tokens": 115,
        "number_input_tokens": 225
    },
...

Lab 6: HPA 구성하기

vCPU 리밋이 걸려있어 수행하기는 어렵습니다.

HAMi v2.7.0부터 AWS Neuron(Trainium/Inferentia)을 정식 지원되어 사용할 수 있지만 vLLM 이 TP=2 구성 환경이라 워크샵환경에서는 활용이 어렵습니다.

다만 이를 위한 구성을 소개하겠습니다.

HPA 동작을 위한 메트릭 서버 구성

# Install metrics server kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml # Wait for deployment kubectl wait --for=condition=available --timeout=300s deployment/metrics-server -n kube-system # Verify installation kubectl top nodes kubectl top pods -n default

KEDA GPU scaler 활용하기

keda-gpu-scaler는 NVIDIA GPU 노드에서 NVML을 직접 읽고 KEDA에 GPU 메트릭을 전달하는 외장 스케일러입니다. DCGM Exporter, Prometheus, PromQL 없이 GPU 사용량과 VRAM 압박을 기준으로 추론 파드를 확장하거나 scale-to-zero로 줄일 수 있습니다.

그림 12 - GPU 노드의 DaemonSet이 NVML 메트릭을 읽고 gRPC External Scaler를 통해 KEDA와 HPA에 전달하는 구조

GPU 노드마다 DaemonSet 파드가 실행됩니다. 이 파드는 libnvidia-ml.so를 통해 GPU의 SM 사용률과 VRAM 사용량을 읽고, gRPC External Scaler protocol로 중앙 KEDA operator에 제공합니다. KEDA는 ScaledObject의 임계값을 기준으로 대상 Deployment의 HPA를 조정합니다.

GPU Node
vLLM Pod → keda-gpu-scaler DaemonSet → KEDA External Scaler → HPA → Deployment
              NVML 메트릭                  gRPC

확인 지표

vllm-inference 프로필은 VRAM 사용률을 기준으로 동작합니다. 기본 target은 80, activation threshold는 5입니다. triton-inference는 GPU 사용률 75를 target으로 사용합니다. 여러 GPU가 있으면 max, min, avg, sum 방식으로 값을 합칠 수 있습니다.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: vllm-inference-scaler
  namespace: default
spec:
  scaleTargetRef:
    name: vllm-deployment
  minReplicaCount: 1
  maxReplicaCount: 10
  triggers:
    - type: external
      metadata:
        scalerAddress: keda-gpu-scaler.keda.svc.cluster.local:6000
        profile: vllm-inference
Neuron 환경

이 Scaler는 NVIDIA NVML과 libnvidia-ml.so를 사용합니다. 이번 실습의 trn1 Neuron 노드에는 바로 적용할 수 없습니다. Neuron은 neuron-monitor, CloudWatch, Prometheus에서 수집한 지표를 기반으로 별도 KEDA external scaler 또는 Prometheus scaler 구성을 만들어야 합니다.

서비스 종료시키기

작업을 마무리한 후, 서비스를 종료시킵니다.

인그레스 지우기

kubectl delete ingress -n monitoring --all
kubectl delete ingress vllm-ingress-simple

vLLM 앱 삭제하기

echo "=== Deleting vLLM application resources ==="

# Delete vLLM specific resources (no ingress or network policies created in this workshop)
kubectl delete service vllm-service --ignore-not-found=true
kubectl delete deployment vllm-deployment --ignore-not-found=true
kubectl delete configmap vllm-shared-config --ignore-not-found=true
kubectl delete pvc s3-model-cache-pvc --ignore-not-found=true
kubectl delete pv s3-model-cache-pv --ignore-not-found=true

echo "vLLM application resources deleted!"

모니터링 스택 제거하기

echo "=== Removing monitoring stack ==="

# Uninstall Helm releases
helm uninstall grafana -n $MONITORING_NAMESPACE --ignore-not-found=true
helm uninstall prometheus -n $MONITORING_NAMESPACE --ignore-not-found=true

# Delete monitoring namespace
kubectl delete namespace $MONITORING_NAMESPACE --ignore-not-found=true

echo "Monitoring stack removed!"

S3 모델 캐시 버킷 삭제하기

echo "=== Deleting S3 model cache bucket ==="

# Get the bucket name
BUCKET_NAME="ai-infra-summit-vllm-models-cache-$(aws sts get-caller-identity --query Account --output text)"

# Delete all objects in the bucket first
aws s3 rm s3://$BUCKET_NAME --recursive --region $AWS_REGION 2>/dev/null || echo "Bucket not found or already empty"

# Delete the bucket
aws s3api delete-bucket --bucket $BUCKET_NAME --region $AWS_REGION 2>/dev/null || echo "Bucket not found or already deleted"

echo "S3 model cache bucket deleted!"

EKS 노드 그룹 삭제하기

삭제 전 워크숍 AWS 콘솔에서 할 일

AWS Cloudformation에 접근 후 노드그룹 스택 종료 방지를 제거한 이후
삭제를 실행시켜주세요.

#
kubectl delete ns ingress-nginx
kubectl delete -n kube-system ds/s3-csi-node
kubectl delete -n kube-system ds/neuron-device-plugin

#
kubectl delete -n kube-system deploy/k8s-neuron-scheduler
kubectl delete -n kube-system deploy/my-scheduler
kubectl delete -n kube-system deploy/s3-csi-controller
kubectl delete -n kube-system deploy/metrics-server

# 노드 그룹 제거
eksctl delete --cluster ai-infra-summit-test-cluster nodegroup neuron-trn1-2x

더 살펴보기

(AWS Workshop) Generative AI on Amazon EKS 에 대해 살펴볼 예정입니다.