[k8s] EKS 운영 트러블슈팅 — Pod 상태 진단, ArgoCD 동기화

Woong·2026년 4월 15일

Docker, k8s

목록 보기
38/38

개요

  • EKS 운영 중 자주 발생하는 문제와 해결 방법을 정리
    • Pod 상태별 대응, 서비스/Ingress 접근 문제, ArgoCD 동기화 실패, 리소스 문제 등

Pod 상태 문제

Pending
  • Pod 이 Pending 상태에서 멈추는 경우
kubectl describe pod <pod_name> -n <namespace>
# Events 섹션 확인
원인메시지 예시해결
노드 부족no nodes available노드 스케일업 또는 리소스 requests 줄이기
nodeSelector 불일치node(s) didn't match selectornodeSelector 값 확인
toleration 누락node(s) had taintstolerations 설정 추가
PVC 바인딩 실패persistentvolumeclaim not foundPVC 상태, StorageClass 확인
  • nodeSelector/toleration 확인 명령어
# 노드 라벨 확인
kubectl get nodes --show-labels | grep nodegroup

# 노드 taint 확인
kubectl describe node <node_name> | grep Taint

CrashLoopBackOff
  • Pod 이 계속 재시작되는 경우
# 현재 로그
kubectl logs <pod_name> -n <namespace>

# 이전 컨테이너 로그 (재시작 전)
kubectl logs <pod_name> -n <namespace> -p
원인해결
앱 에러 (Exception)로그 확인 후 코드 수정
환경변수 누락ConfigMap / Secret 확인
DB 연결 실패DB 접근 가능 여부, 엔드포인트 확인
OOMKilledmemory limits 증가
  • OOMKilled 확인
kubectl get pod <pod_name> -n <namespace> \
  -o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}'

ImagePullBackOff
  • 이미지를 가져오지 못하는 경우
kubectl describe pod <pod_name> -n <namespace> | grep -A5 "Events"
원인해결
이미지 태그 오류kustomization.yaml 의 newTag 확인
ECR 권한 없음노드의 IAM Role 확인
이미지 없음ECR 에 이미지 존재 여부 확인
# ECR 이미지 존재 확인
aws ecr describe-images \
  --repository-name <repo_name> \
  --region ap-northeast-1

서비스 접근 문제

Service 연결 안 됨
# 1. Pod 실행 확인
kubectl get pods -n <namespace> -l app=<app_name>

# 2. Service 확인
kubectl get svc -n <namespace>

# 3. Endpoints 확인 (Pod 와 연결되었는지)
kubectl get endpoints <svc_name> -n <namespace>

# 4. Pod 라벨과 Service selector 일치 확인
kubectl get pod <pod_name> -n <namespace> --show-labels
kubectl get svc <svc_name> -n <namespace> -o yaml | grep selector -A5
  • Endpoints 가 비어있는 경우
    • Pod 라벨과 Service selector 불일치
    • Pod 이 Ready 가 아님 (readinessProbe 실패)
Ingress 접근 안 됨
# Ingress 상태 확인
kubectl get ingress -n <namespace>

# 상세 확인
kubectl describe ingress <ingress_name> -n <namespace>

# ALB 주소 확인
kubectl get ingress <ingress_name> -n <namespace> \
  -o jsonpath='{.status.loadBalancer.ingress[0].hostname}'
  • 주요 원인: path 설정 오류, Service name/port 불일치, namespace 불일치

ArgoCD 동기화 문제

Sync 실패
원인해결
YAML 문법 오류kustomize build 로 로컬 검증
리소스 충돌기존 리소스 삭제 후 재배포
권한 부족ArgoCD ServiceAccount 권한 확인
# 로컬 검증
kustomize build k8s/overlays/dev/<app_name>

# dry-run
kubectl apply -k k8s/overlays/dev/<app_name> --dry-run=client
이미지 태그 업데이트 안 됨
  • CI/CD 후에도 이전 이미지로 실행되는 경우
# 현재 실행 중인 이미지 확인
kubectl get pod <pod_name> -n <namespace> \
  -o jsonpath='{.spec.containers[0].image}'

# kustomization.yaml 의 newTag 확인
cat k8s/overlays/dev/<app_name>/kustomization.yaml | grep newTag
  • 해결 순서
    1. kustomization.yaml 의 newTag 값 확인
    2. ArgoCD 에서 Refresh → Sync
    3. 그래도 안 되면 Pod 삭제하여 재생성

리소스 문제

OOMKilled
kubectl describe pod <pod_name> -n <namespace> | grep -i oom
  • 해결: patch-deployment.yaml 에서 memory limits 증가
resources:
  limits:
    memory: 1Gi    # 기존 512Mi → 1Gi
CPU Throttling
  • 앱 응답이 느려지거나 타임아웃 발생
  • Grafana 대시보드에서 CPU throttling 확인
resources:
  requests:
    cpu: 200m      # 증가
  limits:
    cpu: 1000m     # 증가

디버깅 명령어 모음

# Pod 상태 한눈에 보기
kubectl get pods -n <namespace> -o wide

# Running 아닌 Pod 찾기
kubectl get pods -n <namespace> | grep -v Running

# 최근 이벤트 확인
kubectl get events -n <namespace> --sort-by='.lastTimestamp' | tail -20

# 특정 앱 전체 리소스 확인
kubectl get all -n <namespace> -l app=<app_name>

# Pod 내부 접속
kubectl exec -it <pod_name> -n <namespace> -- /bin/sh

reference

0개의 댓글