개요
- EKS 운영 중 자주 발생하는 문제와 해결 방법을 정리
- Pod 상태별 대응, 서비스/Ingress 접근 문제, ArgoCD 동기화 실패, 리소스 문제 등
Pod 상태 문제
Pending
- Pod 이
Pending 상태에서 멈추는 경우
kubectl describe pod <pod_name> -n <namespace>
| 원인 | 메시지 예시 | 해결 |
|---|
| 노드 부족 | no nodes available | 노드 스케일업 또는 리소스 requests 줄이기 |
| nodeSelector 불일치 | node(s) didn't match selector | nodeSelector 값 확인 |
| toleration 누락 | node(s) had taints | tolerations 설정 추가 |
| PVC 바인딩 실패 | persistentvolumeclaim not found | PVC 상태, StorageClass 확인 |
- nodeSelector/toleration 확인 명령어
kubectl get nodes --show-labels | grep nodegroup
kubectl describe node <node_name> | grep Taint
CrashLoopBackOff
kubectl logs <pod_name> -n <namespace>
kubectl logs <pod_name> -n <namespace> -p
| 원인 | 해결 |
|---|
| 앱 에러 (Exception) | 로그 확인 후 코드 수정 |
| 환경변수 누락 | ConfigMap / Secret 확인 |
| DB 연결 실패 | DB 접근 가능 여부, 엔드포인트 확인 |
| OOMKilled | memory limits 증가 |
kubectl get pod <pod_name> -n <namespace> \
-o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}'
ImagePullBackOff
kubectl describe pod <pod_name> -n <namespace> | grep -A5 "Events"
| 원인 | 해결 |
|---|
| 이미지 태그 오류 | kustomization.yaml 의 newTag 확인 |
| ECR 권한 없음 | 노드의 IAM Role 확인 |
| 이미지 없음 | ECR 에 이미지 존재 여부 확인 |
aws ecr describe-images \
--repository-name <repo_name> \
--region ap-northeast-1
서비스 접근 문제
Service 연결 안 됨
kubectl get pods -n <namespace> -l app=<app_name>
kubectl get svc -n <namespace>
kubectl get endpoints <svc_name> -n <namespace>
kubectl get pod <pod_name> -n <namespace> --show-labels
kubectl get svc <svc_name> -n <namespace> -o yaml | grep selector -A5
- Endpoints 가 비어있는 경우
- Pod 라벨과 Service selector 불일치
- Pod 이 Ready 가 아님 (readinessProbe 실패)
Ingress 접근 안 됨
kubectl get ingress -n <namespace>
kubectl describe ingress <ingress_name> -n <namespace>
kubectl get ingress <ingress_name> -n <namespace> \
-o jsonpath='{.status.loadBalancer.ingress[0].hostname}'
- 주요 원인: path 설정 오류, Service name/port 불일치, namespace 불일치
ArgoCD 동기화 문제
Sync 실패
| 원인 | 해결 |
|---|
| YAML 문법 오류 | kustomize build 로 로컬 검증 |
| 리소스 충돌 | 기존 리소스 삭제 후 재배포 |
| 권한 부족 | ArgoCD ServiceAccount 권한 확인 |
kustomize build k8s/overlays/dev/<app_name>
kubectl apply -k k8s/overlays/dev/<app_name> --dry-run=client
이미지 태그 업데이트 안 됨
- CI/CD 후에도 이전 이미지로 실행되는 경우
kubectl get pod <pod_name> -n <namespace> \
-o jsonpath='{.spec.containers[0].image}'
cat k8s/overlays/dev/<app_name>/kustomization.yaml | grep newTag
- 해결 순서
kustomization.yaml 의 newTag 값 확인
- ArgoCD 에서 Refresh → Sync
- 그래도 안 되면 Pod 삭제하여 재생성
리소스 문제
OOMKilled
kubectl describe pod <pod_name> -n <namespace> | grep -i oom
- 해결:
patch-deployment.yaml 에서 memory limits 증가
resources:
limits:
memory: 1Gi
CPU Throttling
- 앱 응답이 느려지거나 타임아웃 발생
- Grafana 대시보드에서 CPU throttling 확인
resources:
requests:
cpu: 200m
limits:
cpu: 1000m
디버깅 명령어 모음
kubectl get pods -n <namespace> -o wide
kubectl get pods -n <namespace> | grep -v Running
kubectl get events -n <namespace> --sort-by='.lastTimestamp' | tail -20
kubectl get all -n <namespace> -l app=<app_name>
kubectl exec -it <pod_name> -n <namespace> -- /bin/sh
reference