华为与高通达成多年期专利交叉许可协议
Readhub
华为与高通达成覆盖 5G、计算、人工智能和网络等技术领域的多年期专利交叉许可协议,高通还将购买华为在计算、人工智能、网络及其他技术领域的部分美国专利。
Балл: 57.18Уверенность: 54%
ПодробнееЗагружаем каталог…
НАВИГАТОР ПО ВОЗМОЖНОСТЯМ ИИ
Найдите свой ИИ-инструмент. Бесплатный доступ, пробные периоды и кредиты — в одном месте.
Readhub
华为与高通达成覆盖 5G、计算、人工智能和网络等技术领域的多年期专利交叉许可协议,高通还将购买华为在计算、人工智能、网络及其他技术领域的部分美国专利。
Балл: 57.18Уверенность: 54%
Подробнее掘金
Dart 3 可用 Record 的 `wait` 并发等待不同类型的 Future,既保留具体返回类型,也省去索引取值和手动类型转换。
Балл: 57.15Уверенность: 54%
Подробнееvelog
무슨 일이 있었나 오픈AI가 2026년 10월 6일(현지시간), ChatGPT가 생성한 텍스트에 워터마크를 삽입하는 기능을 EU 지역 사용자에 한해 기본으로 적용하겠다고 공식 발표했다. 오픈AI는 자사 공식 발표문에서 워터마크가 어떤 범위에 적용되는지, 탐지는 어떻게 이뤄지는지, 그리고 왜 탐지 권한을 일단 연구자 등 제한된 그룹에만 열어두는지를 설명했다. 이번 조치는 EU AI법(AI Act) 50조에 따른 것이다. 해당 조항은 생성형 AI 사업자가 AI로 만든 콘텐츠를 기계가 판독 가능한 형태로 표시하도록 의무화하고 있으며, 2026년 8월 2일부로 발효됐다. 다만 이미 시장에 나와 있던 서비스에는 2026년 12월 2일까지 유예기간이 주어진 상태다. 경쟁사인 앤트로픽은 이미 지난 8월부터 새로 출시하는 클로드 모델에 텍스트를 복사하거나 일부 편집해도 사라지지 않는 비가시 워터마크를 심어왔다. 반면 오픈AI는 2024년 한 차례 자체 연구에서 ChatGPT 이용자의 약 3분의 1이 워터마크 도입 시 사용을 줄이겠다고 답한 결과가 나오자 관련 계획을 보류한 바 있다. 이번 발표로 오픈AI도 결국 규제를 계기로 워터마크 도입에 나선 모양새다. 탐지 기능은 아직 전면 공개되지 않았으며, 규제당국·수사기관·언론·팩트체커·연구기관·EU 컴플라이언스 기업 등에 한해 비공개 프리뷰 형태로 제공된다. 왜 중요한가 텍스트 워터마킹은 이미지·음성 워터마킹과 달리 기술적 난도가 높다. 모델이 다음 단어를 고를 때 특정 단어를 통계적으로 더 자주 선택하도록 미세하게 편향을 주는 방식으로 구현되는데, 사람 눈에는 티가 나지 않지만 충분히 긴 글에서는 탐지기가 패턴을 잡아낼 수 있다. 문제는 이 패턴이 짧은 글이나 번역·편집을 거친 글에서는 쉽게 깨질 수 있어 완벽한 해법이 아니라는 점이다. 그럼에도 오픈AI의 이번 결정은 업계 전반에 신호를 준다. 세계 최대 AI 서비스 사업자가 규제가 없는 지역에서는 워터마크를 적용하지 않기로 한 것은, 결국 '어느 시장이 규제하느냐'가 AI 투명성 기능의 실질적 적용 범위를 좌우한다는 사실을 보여준다. 동시에 생성형 AI를 활용한 허위정보·표절·학술 부정행위 문제가 커지는 가운데, 콘텐츠 출처 판별 기술이 일부 기관에만 열려 있는 '비대칭 탐지' 구조가 당분간 이어질 전망이어서, 워터마크의 실효성과 공정한 접근성을 둘러싼 논쟁은 계속될 것으로 보인다. 원문: https://openai.com/index/eu-text-provenance
Балл: 54.39Уверенность: 49%
Подробнееvelog
Travelers searching for something more adventurous can add an off-road experience to their Bali holiday. A Bali UTV ride offers an opportunity to explore selected outdoor routes while enjoying the island's natural surroundings. The experience can take participants away from busy tourist roads and into areas featuring village landscapes, tropical plants, open fields, and other scenery that makes the journey feel more connected to the outdoors. The fun of an off-road activity comes from not knowing exactly what type of terrain will appear around the next section. Depending on the route, participants may travel across dusty tracks, uneven ground, muddy patches, or other surfaces that make the ride more engaging. For a different setting, Buggy Bali Beach offers an opportunity to enjoy recreational driving near the coast. The combination of an off-road vehicle and beach surroundings creates a distinct experience for visitors who want to spend time outdoors. Proper preparation can help participants feel more comfortable throughout the adventure. Before departure, the operator normally explains how the vehicle works and provides essential information about safe riding. Participants should understand how to control the vehicle, follow the designated route, and respond to instructions from guides. It is also wise to check the activity requirements before booking, especially details concerning age limits, passenger arrangements, safety equipment, route length, and weather-related changes. A UTV adventure can bring a more dynamic element to a Bali itinerary without taking away from the island's relaxing atmosphere. It gives travelers an opportunity to combine sightseeing with an activity that involves movement, nature, and exploration. Whether you are looking for an exciting experience with friends or want to add variety to a couple or family trip, an off-road UTV ride can offer a memorable way to enjoy Bali beyond its conventional tourist experiences.
Балл: 54.39Уверенность: 49%
Подробнееvelog
Navigating the Aesthetic Landscape in Modern Delhi Choosing to undergo aesthetic or reconstructive transformation represents a meaningful personal milestone. Today, prospective patients exploring Cosmetic Surgery in Delhi encounter a dynamic healthcare landscape filled with diverse procedural techniques, advanced medical technologies, and varying philosophies of care. Navigating this vast array of information requires a structured, patient-centric approach grounded in scientific clarity, safety verification, and biological readiness. Modern aesthetic medicine has transitioned away from exaggerated physical alterations toward subtle, bio-harmonious enhancements. Rather than altering baseline anatomical features beyond recognition, contemporary surgical principles emphasize preserving microvascular supply, respecting natural tissue planes, and enhancing underlying facial and body symmetry. At Vitality Aesthetics in Vasant Vihar, South Delhi, this patient journey is guided with clinical rigor and individualized attention. Under the healthcare vision of Rachita Kolli, DNHE (Certified Nutritionist & Healthcare Consultant), the practice combines refined procedural execution with pre- and post-operative metabolic conditioning, providing patients with an empowering, safety-first framework for their aesthetic goals. Defining Your Motivations and Establishing Realistic Goals The initial phase of any surgical journey begins with self-reflection and establishing clear clinical objectives. Patients seek aesthetic refinements for many reasons, including age-related tissue redrapery, post-pregnancy body restoration, corrective scar refinement, or inherent anatomical asymmetry. Clarifying these personal triggers helps align expectations with what surgical interventions can realistically accomplish. When evaluating your goals with a Cosmetic Plastic Surgeon in Delhi, focus on structural harmony rather than perfection. Surgical procedures alter soft tissue, reposition deep fascial layers, or adjust contours within the boundaries of your unique skeletal architecture and dermal elasticity. Understanding that elective procedures aim to enhance natural features rather than duplicate idealized templates ensures higher satisfaction and a balanced psychological outcome post-procedure. The Pre-Consultation Audit: Preparing for Your Initial Visit Preparing thoroughly for your initial clinical visit maximizes the value of your consultation. A informative pre-consultation audit allows your medical team to evaluate your candidacy with complete accuracy: Documenting Medical History: Compile a complete record of past surgical procedures, existing medical conditions (such as hypertension, thyroid disorders, or glycemic fluctuations), known drug allergies, and active prescription regimens. Listing Supplements and Over-the-Counter Products: Identify all vitamins, herbal extracts, anti-inflammatory agents, or blood-thinning supplements consumed regularly, as many interfere with normal coagulation. Formulating Specific Questions: Write down key inquiries regarding the surgeon's qualifications, facility accreditation, anesthesia options, surgical vectors, and expected recovery timelines. Clarifying Lifestyle Factors: Be prepared to discuss dietary habits, physical activity levels, alcohol consumption, and active or past nicotine exposure. Entering a Plastic Surgery Clinic in Delhi with organized information enables a transparent, efficient dialogue regarding your suitability for surgery. Comprehensive Physiological Evaluation and Risk Stratification Patient safety begins long before entering an operating room. A rigorous pre-operative assessment verifies that your cardiovascular, metabolic, and tissue dynamics are fully prepared to support anesthesia and healing. When consulting a Best Aesthetic Surgeon in Delhi, expect a systematic health evaluation including: Comprehensive Hematological Panels: Complete blood counts, liver and kidney performance tests, blood glucose stabilization assessments, and coagulation profiles to prevent operative bleeding risks. Cardiovascular and Pulmonary Screenings: Baseline electrocardiograms (ECG) and chest imaging when indicated to confirm systemic stamina under anesthesia. Dermal and Structural Mapping: Objective physical examination assessing skin elasticity, subcutaneous fat distribution, vascular integrity, and underlying muscle tone. Microvascular and Perfusion Audits: Evaluating micro-circulation and stopping all nicotine exposure weeks prior to surgery, as nicotine constricts capillaries and significantly delays wound edge closure. Patients searching for Plastic & Cosmetic Surgery Near Me benefit directly from this detailed medical check, which lowers procedural risks and establishes a solid foundation for repair. Core Surgical Categories and Anatomical Nuances Contemporary aesthetic care encompasses specialized disciplines tailored to different body regions and structural layers. Understanding these core categories helps patients understand how specific procedural techniques address distinct anatomical concerns: Facial Contouring and Rejuvenation: Procedures such as blepharoplasty, rhytidectomy, neck lifts, and rhinoplasty restore youthful proportions by repositioning deep muscular SMAS layers, redraping skin without tension, and correcting micro-vascular volume depletion through target lipotransfer during facial rejuvenation. Body Structure and Silhouette Sculpting: Interventions including multi-plane liposuction, abdominoplasty, and body lifts repair separated abdominal fascia, remove recalcitrant adipose compartments, and restore firm physical boundaries during custom body contouring. Breast Harmonization and Reconstruction: Augmentation, mastopexy, reduction, or reconstructive procedures utilize precise anatomical mapping to optimize symmetry, preserve glandular architecture, and ensure stable long-term tissue support. Targeted Tissue Optimization: Reconstructive scar revisions and delicate soft tissue repairs focus on tension-free wound closure to restore mobility and refine surface texture. Consulting a Best Plastic Surgeon in Delhi ensures that procedural choices are customized directly to your tissue characteristics and overall health profile. Integrating Cellular Nutrition into the Surgical Journey Surgical skill constitutes one half of a successful visual result; the body's internal biological capacity to repair tissue constitutes the other. Surgical incisions create controlled tissue stress, triggering complex cellular repair phases including inflammation, micro-vascular proliferation, and collagen matrix deposition. How effectively these phases progress depends heavily on baseline cellular nutrition. At Vitality Aesthetics, healthcare leadership under Rachita Kolli, DNHE , incorporates specialized metabolic preparation directly into the care plan: Pre-Surgical Biological Fortification: Correcting sub-clinical micronutrient deficiencies, building essential amino acid reserves, and optimizing cellular antioxidant capacity to prepare tissues for operative stress. Systemic Inflammatory Control: Employing targeted clinical nutrition to manage post-operative swelling, accelerate hematoma resorption, and stabilize oxidative markers. Accelerated Collagen Maturation: Delivering vital mineral cofactors, bioflavonoids, and bioavailable proteins required for strong incisional healing and clean scar formation. Patients choosing a Best Aesthetic Clinic in Delhi experience how aligning surgical intervention with cellular health supports smoother recovery phases and clean structural outcomes. Evaluating Clinical Infrastructure, Safety Standards, and Sterile Care The clinical setting where your procedure takes place is as critical as the surgeon’s technical skill. Sterile operating environments minimize infection risks and safeguard patient well-being throughout perioperative care. Key safety protocols to verify at a Top Plastic Surgery Clinic South Delhi include: HEPA-Filtered Operating Environments: Surgical suites must feature positive-pressure airflow systems equipped with High-Efficiency Particulate Air (HEPA) filtration to maintain pristine cleanroom standards. Rigorous Sterilization Workflows: Multi-tiered autoclaving, chemical verification indicators, and single-use components protect against cross-contamination. Board-Certified Anesthesiology Support: Continuous real-time tracking of vital signs, cardiac rhythms, fluid management, and airway safety administered by dedicated senior anesthesiologists. Comprehensive Emergency Readiness: Advanced cardiac life support infrastructure, emergency resuscitation kits, and formal transfer arrangements with tertiary medical centers ensure complete security. Selecting an accredited Cosmetic Surgery Clinic in Delhi gives patients complete confidence regarding facility safety. Understanding the Financial Dynamics and Cost Structure Navigating the financial aspects of cosmetic surgery requires understanding the comprehensive clinical investment involved in safe, individualized medical care. Ethical medical practices avoid flat-rate or standardized pricing because every procedural design is tailored to the patient's biological complexity and scope of treatment. Overall investment varies according to several key clinical parameters: Procedural Extent and Complexity: The total operating theater duration, technical difficulty, and whether isolated or combined procedures are performed. Specialist Qualifications and Expertise: The surgeon's subspecialty training, board certifications, clinical experience, and recognized track record. Facility and Anesthesia Overhead: The utilization of accredited cleanroom suites, specialized surgical monitoring equipment, and board-certified anesthesiologist care. Comprehensive Perioperative Support: Pre-operative diagnostic lab panels, custom compression garments, routine post-surgical follow-ups, and integrated metabolic nutrition plans.
velog
🎯 2026년 목표 분야 목표 ☁️ Cloud 클라우드 환경 운영 경험 쌓기 (상세 내용 추후 작성) 🔐 Security 보안관제 자동화 경험 ✍️ Blog 블로그 꾸준히 작성 📚 자격증 정보처리기사 / 정보보안기사 / AWS-SAA 🏆 장기 자격증 CISA / CISSP (정보보안 관련 경력 요건 충족 후 도전) 💻 Development 개발 공부 꾸준히 진행 🇺🇸 English 영어 공부 🏃 Exercise 꾸준한 운동 🔄 Review 리팩토링 과정 복습 및 Velog 재작성 ☁️ Cloud Security 클라우드 보안 복습 및 Velog 작성 🎓 Next Year 석사 과정 예정 🗺️ 공부 로드맵 - Version 10월 🟢 완료 | 🟡 진행중 | 🔴 미완료 | ⚪ 시작 전 순서 공부 상태 / 목표 진행률 1 🐍 파이썬 🟢 점프 투 파이썬 완강 → 지속적인 복습 + 개인 프로젝트 100% 2 📘 정보처리기사 실기 🟡 10월 25일 시험 / 합격 목표 50% 3 🔐 모의해킹 - 노말틱 ⚪ 정보처리기사 실기 끝나고 11월 시작 0% 4 🇺🇸 영어 ⚪ 11월부터 시작 0% 5 🌐 혼자 공부하는 네트워크 🟢 완강 → 지속적인 복습 100% 6 🔄 리팩토링 과정 ⚪ 복습 + 실습 + 블로그 재작성 0% 7 🐳 Docker ⚪ 모의해킹 완료 후 시작 0% 8 ☁️ AWS 보안 가이드 ⚪ Docker 강의 완료 후 시작 0% 9 ☕ JAVA ⚪ 파이썬 활용 능력이 어느 정도 자리 잡은 후 시작 0% 📅 금주 목표 - 10/05(월) ~ 10/11(일) 🟢 완료 | 🟡 진행중 | 🔴 미완료 | ⚪ 시작 전 분야 목표 진행 상황 진행률 🌐 Network 혼자 공부하는 네트워크 복습 🟡 30% 🐍 Python Project Nmap과 유사한 기능 직접 구현 🟡 40% 📘 정보처리기사 개념 정리 🟡 40% 📘 정보처리기사 2020년도 기출 3회차 풀이 ⚪ 0% 📘 정보처리기사 2021년도 기출 2회차 풀이 ⚪ 0% 📘 정보처리기사 2022년도 기출 2회차 풀이 ⚪ 0% 📘 정보처리기사 2023년도 기출 1회차 풀이 ⚪ 0% 📘 정보처리기사 2024년도 기출 1회차 풀이 ⚪ 0% 📘 정보처리기사 2025년도 기출 1회차 풀이 ⚪ 0% 📘 정보처리기사 2026년도 기출 1회차 풀이 ⚪ 0% ✅ 오늘의 목표 🟢 완료 | 🟡 진행중 | 🔴 미완료 | ⚪ 시작 전 상태 할 일 ⚪ 💰 Household Account 작성 - 23:50 ⚪ 📘 **정보처리기사 개념 정리 + 블로그 정리 ⚪ 📘 정보처리기사 DAY 6 - 문제풀이 ⚪ 🌐 혼자 공부하는 네트워크 복습 DAY 6 ⚪ 🐍 Python Security Scanner DAY 6 - Nmap 기능 구현 + 블로그 정리 💻 오늘 완료한 일 상태 완료 내역 ✅ ✅ ✅ ✅ ✅ ✅
Балл: 54.37Уверенность: 49%
Подробнееvelog
지금 구조라면 GPU node를 기존 Compute Cluster에 그냥 추가하기보다는, GPU 전용 Kubernetes Cluster를 하나 더 만들고 기존 Compute / Storage / Admin과 역할을 분리 하는 안을 우선 추천합니다. 즉 최종적으로는 4-cluster architecture 입니다. ┌──────────────────────────────┐ │ Admin Cluster │ │ │ │ CI/CD / GitOps │ │ Keycloak / SSO │ │ Cluster Management │ │ Registry / Helm Repo │ │ Monitoring / Logging(*) │ └──────────────┬───────────────┘ │ Management / GitOps / API │ ┌───────────────────────────┼───────────────────────────┐ │ │ │ ▼ ▼ ▼ ┌────────────────────┐ ┌────────────────────┐ ┌────────────────────┐ │ Compute Cluster │ │ GPU Cluster │ │ Storage Cluster │ │ │ │ │ │ │ │ CPU workload │ │ B300 GPU nodes │ │ AIStor │ │ API / Backend │ │ vLLM / inference │ │ MinIO/S3 │ │ preprocessing │ │ training │ │ durable dataset │ │ application │ │ OCR/VLM(*) │ │ model/checkpoint │ └─────────┬──────────┘ └─────────┬──────────┘ └─────────┬──────────┘ │ │ │ └──────────────────────────┼───────────────────────────┘ │ Cilium ClusterMesh BGP / Native Routing │ Ethernet L3 Fabric 여기서 중요한 것은 ClusterMesh는 service/workload connectivity를 제공하는 계층이고, AIStor의 대용량 S3 data path나 GPU의 NDR/IB fabric까지 무조건 ClusterMesh에 태우자는 의미는 아닙니다. 1. 각 Cluster 역할 저라면 경계를 다음처럼 잡겠습니다. Cluster Node 핵심 역할 넣지 않을 것 Admin 일반 CPU GitOps, CI/CD, Keycloak, Registry, cluster lifecycle, observability AI workload Compute CPU Compute API/backend, preprocessing, orchestration, 일반 application 대형 GPU serving GPU B300 vLLM, LLM/VLM inference, training, GPU batch AIStor Storage Storage node AIStor/S3, model/dataset/checkpoint application/GPU workload 현재 Compute와 Storage가 이미 ClusterMesh라면: 현재 Compute Cluster │ │ Cilium ClusterMesh │ Storage Cluster 여기에 GPU cluster를: Compute / \ / \ / \ GPU ───── Storage ClusterMesh 로 참여시키는 것입니다. Admin cluster까지 반드시 동일한 ClusterMesh에 넣을 필요는 없습니다. 오히려 저는 Admin은 management plane으로 분리 하는 쪽을 선호합니다. 2. Admin Cluster는 ClusterMesh와 분리하는 게 좋다 Admin cluster 역할을 보면: Admin ├── Argo CD / GitOps ├── CI/CD ├── Keycloak ├── Registry ├── Helm repository ├── Cluster management └── Monitoring 이므로 application data plane과 성격이 다릅니다. 따라서: Admin Cluster │ Management Network │ ┌───────────────┼──────────────┐ ▼ ▼ ▼ Compute GPU Storage 로 관리하고, ClusterMesh는: ┌──── Compute ────┐ │ │ │ ClusterMesh │ │ │ GPU ───────── Storage 정도로 제한하는 것이 좋습니다. 이렇게 하면 GPU/Compute/Storage data plane에 문제가 발생해도 관리 plane의 blast radius를 줄일 수 있습니다. 3. GPU Cluster를 별도로 만드는 이유 기존 Compute Cluster에 GPU node를 추가하는 것도 기술적으로 가능합니다. Compute Cluster CPU node CPU node CPU node GPU B300 GPU B300 GPU B300 하지만 지금 정도 규모/구조라면 추천하지 않습니다. GPU node에는 다음처럼 일반 Compute와 상당히 다른 node-level 설정이 들어가기 때문입니다. GPU Node RHEL 10 │ ├─ NVIDIA Driver ├─ Container Toolkit ├─ GPU Operator ├─ NVIDIA Device Plugin ├─ DCGM ├─ NFD │ ├─ CPU Manager = static ├─ Topology Manager = restricted │ ├─ HugePages(optional) │ ├─ NVLink / NVSwitch ├─ NCCL │ ├─ NDR InfiniBand ├─ RDMA │ └─ NVMe x8 local cache 운영 lifecycle 자체가 CPU compute node와 다릅니다. 특히 GPU driver/CUDA/NCCL/GPU Operator upgrade 때문에 GPU cluster의 maintenance cycle을 Compute와 독립시키는 가치가 큽니다. 4. GPU Cluster 내부 구성 저라면 GPU cluster는 최소 다음과 같이 만듭니다. GPU Kubernetes Cluster ┌────────────────────────────┐ │ Control Plane x3 │ │ │ │ 일반 CPU server/VM │ │ GPU 없음 │ └──────────────┬─────────────┘ │ Kubernetes API │ ┌───────────────┼─────────────────┐ │ │ │ ▼ ▼ ▼ GPU Node 01 GPU Node 02 GPU Node N B300 x8 B300 x8 B300 x8 NVMe x8 NVMe x8 NVMe x8 RAID0 RAID0 RAID0 XFS XFS XFS /model-cache /model-cache /model-cache Control Plane을 B300 node 위에 올리는 것은 피합니다. 5. GPU node network는 최소 3가지 성격으로 분리 여기가 특히 중요합니다. GPU Node │ ┌──────────────────┼───────────────────┐ │ │ │ ▼ ▼ ▼ Management Ethernet Data NDR IB Network Network Fabric K8s API Cilium/BGP NCCL SSH/OOB ClusterMesh GPU↔GPU monitoring AIStor/S3 RDMA inference API 논리적으로: Ethernet GPU │ ├── Pod Network ├── Service Network ├── Cilium ├── BGP ├── ClusterMesh │ ├── Compute ↔ GPU └── GPU ↔ AIStor 현재 사용하는 Cilium native routing + BGP + ECMP 를 이쪽에 그대로 확장하는 것이 자연스럽습니다. InfiniBand/NDR GPU Node 1 │ NDR │ GPU Node 2 │ NDR │ GPU Node N 여기는 ClusterMesh traffic용이 아니라 우선: NCCL MPI GPU collective RDMA 전용으로 봅니다. 6. GPU Pod에서는 multi-network가 필요해질 수 있다 일반 serving Pod: vLLM Pod │ eth0 │ Cilium │ Compute/API 만 있으면 됩니다. 하지만 multi-node GPU workload라면: GPU Pod ┌─────┴─────┐ │ │ eth0 ib0 │ │ Cilium NDR │ │ API/S3 NCCL/RDMA 형태가 좋습니다. 따라서 GPU cluster에는 Multus + NVIDIA Network Operator / RDMA device plugin 계열 을 검토할 가치가 있습니다. 즉: Primary CNI ──────────── Cilium │ └─ application/service traffic Secondary Network ───────────────── Multus │ └─ IB/RDMA │ └─ NCCL 입니다. Cilium을 없애고 IB를 쓰는 게 아니라 용도가 다른 두 network를 Pod에 제공 하는 구조입니다. 7. Compute → GPU inference path 예를 들어 Compute cluster의 backend가 GPU cluster의 vLLM을 호출한다면: Compute Cluster Application Pod │ │ HTTP/gRPC ▼ Cilium │ │ ClusterMesh ▼ GPU Cluster │ ▼ Inference Service │ ▼ vLLM Pod │ ▼ B300 가 됩니다. ClusterMesh를 이미 운영 중이라면 이 구조가 자연스럽습니다. 서비스 이름도 global service/discovery 전략을 정해서: llm-serving.gpu... 같은 논리 endpoint로 노출할 수 있습니다. 다만 모든 GPU Pod를 Compute에서 직접 접근시키기보다 GPU cluster에 Gateway/API 계층을 두는 것 을 권합니다. Compute │ ▼ GPU Gateway / LB │ ├── DeepSeek service ├── GLM service ├── Qwen service └── OCR service │ ▼ GPU Pod 이렇게 해야 backend가 GPU topology/Pod 배치를 몰라도 됩니다. 8. GPU → Storage path는 조금 다르게 봐야 한다 여기서 중요한 설계 포인트가 있습니다. 논리적으로는: GPU Pod │ ClusterMesh │ AIStor Service 라고 만들 수 있습니다. 하지만 대규모 model/dataset/checkpoint S3 traffic까지 ClusterMesh service path에 의존할 필요는 없습니다. 저라면: GPU Cluster │ │ Ethernet L3 │ │ S3/HTTPS ▼ AIStor VIP / LB │ Storage Cluster 형태의 직접 routed data path 를 우선 검토합니다. 즉: Small/control/service traffic Compute ←── ClusterMesh ──→ GPU ↘ Storage Bulk storage traffic GPU =======================> AIStor L3 Ethernet / S3 입니다. Cilium/BGP를 사용하고 있으므로 underlay routing을 잘 설계하면 이 방식과 궁합이 좋습니다. 9. Storage Cluster도 역할을 매우 좁게 유지 Storage cluster는: Storage Kubernetes Cluster Control Plane x3 AIStor Node ├ NVMe ├ NVMe ├ ... └ high-speed Ethernet AIStor Node ├ NVMe └ ... AIStor Node └ ... 로 유지합니다. 여기에 GPU workload나 일반 application을 넣지 않습니다. 그리고 bucket 역할은: AIStor s3://models/ s3://datasets/ s3://checkpoints/ s3://artifacts/ 정도로 분리합니다. 10. Local NVMe와 AIStor 관계 GPU node에는 앞에서 정한: NVMe x8 │ RAID0 │ XFS │ /model-cache 를 둡니다. 전체 storage hierarchy는: AIStor durable / source │ │ S3 ▼ GPU Local NVMe 8× RAID0 + XFS cache/stage │ ▼ CPU RAM │ ▼ GPU HBM │ ┌────┴────┐ │ │ Weight KV 가 됩니다. AIStor = source of truth , local NVMe는 언제든 재생성 가능한 cache라는 원칙을 유지합니다. 11. Admin Cluster에서는 GitOps로 네 cluster를 통제 Air-gap이므로 이 구조가 특히 중요합니다. Admin Cluster │ ┌─────────────────┼─────────────────┐ │ │ │ Git Repo OCI Registry Keycloak │ │ │ └─────────────────┼─────────────────┘ │ GitOps │ ┌─────────────────┼──────────────────┐ ▼ ▼ ▼ Compute GPU Storage Private registry에는: NVIDIA GPU Operator NVIDIA Network Operator NVIDIA driver container DCGM exporter vLLM NCCL Cilium AIStor application images 등 필요한 모든 artifact를 mirror합니다. GPU cluster가 인터넷을 직접 볼 필요가 없도록 합니다. 12. Observability도 Admin 쪽으로 집계 각 cluster에는 exporter/agent를 두고: Compute Prometheus agent ─┐ GPU DCGM exporter ────┼──> Admin Observability Prometheus agent ─┤ Storage metrics agent ────┘ 형태로 중앙 집계하는 게 좋습니다. GPU에서는 특히: GPU utilization HBM usage GPU temperature/power NVLink traffic PCIe traffic ECC XID error GPU throttling CPU NUMA local NVMe throughput NCCL IB/RDMA counters model latency TTFT TPOT tokens/sec KV-cache utilization 까지 같이 봐야 합니다. 그래야 GPU utilization이 떨어졌을 때: AIStor? ↓ Ethernet? ↓ NVMe? ↓ CPU/DataLoader? ↓ GPU? ↓ NVLink/NCCL? 중 어디가 병목인지 찾을 수 있습니다. 13. 최종적으로 제가 권하는 전체 구조 현재 환경을 크게 뜯어고치지 않고 GPU를 추가한다면 이 형태가 가장 자연스럽습니다. ADMIN CLUSTER ┌──────────────────────────┐ │ CI/CD / GitOps │ │ Keycloak │ │ Registry / Helm │ │ Cluster Management │ │ Observability │ └────────────┬─────────────┘ │ Management Plane │ ┌─────────────────────┼─────────────────────┐ │ │ │ ▼ ▼ ▼ ┌───────────────┐ ┌────────────────┐ ┌───────────────┐ │ COMPUTE │ │ GPU CLUSTER │ │ STORAGE │ │ CLUSTER │ │ │ │ CLUSTER │ │ │ │ B300 x8/node │ │ │ │ CPU workload │ │ │ │ AIStor │ │ backend │ │ vLLM │ │ S3 │ │ API │ │ LLM/VLM │ │ model │ │ preprocess │ │ training │ │ dataset │ │ │ │ │ │ checkpoint │ └───────┬───────┘ └───────┬────────┘ └───────┬───────┘ │ │ │ │ Cilium ClusterMesh │ └─────────────────────┼──────────────────────┘ │ Ethernet Fabric BGP + ECMP + L3 │ ┌─────────────┴────────────┐ │ │ Service traffic S3 bulk │ GPU ↔ AIStor Separate NDR IB Fabric GPU Node ───────── GPU Node │ │ └──── NDR IB ─────┘ NCCL/RDMA 여기서 제가 특히 설계 원칙으로 고정할 것 은 Admin=Management , Compute=CPU application , GPU=accelerated workload , Storage=durable object storage 의 4-way separation입니다. 그리고 Cilium ClusterMesh는 Compute/GPU/Storage의 서비스 연결 , Ethernet L3 fabric은 AIStor 같은 bulk data path, NDR IB는 GPU collective/RDMA 라는 식으로 네트워크 역할도 분리합니다. 추가로 GPU cluster kubelet은 앞에서 이야기한 CPU Manager=static , Topology Manager=restricted , scope=pod 를 baseline으로 하고, GPU Operator/NFD/DCGM 및 필요 시 Network Operator+Multus/RDMA를 얹는 구조가 좋습니다. 이 정도면 이후 B300 증설이나 GPU driver/CUDA 업그레이드가 기존 Compute/Storage cluster에 미치는 영향도 상당히 줄일 수 있습니다. === 가능합니다. 지금 요구사항이면 GPU Cluster 안에 “Inference Platform” 계층을 하나 만드는 것 이 좋습니다. 사용자는 개별 vLLM Pod를 직접 호출하지 않고, 항상 Inference Gateway 를 통해 접근하게 합니다. 핵심 구조는 아래처럼 잡겠습니다. ┌──────────────────────────────┐ │ Admin Cluster │ │ │ │ Keycloak / GitOps / CI/CD │ │ Model Schedule / Policy │ │ Observability / Registry │ └──────────────┬───────────────┘ │ configuration / identity │ ▼ ┌──────────────────────────────── GPU Cluster ───────────────────────────────┐ │ │ │ Users / Compute Cluster │ │ │ │ │ ▼ │ │ ┌─────────────────────┐ │ │ │ Inference Gateway │ │ │ │ │ │ │ │ OIDC / Keycloak │ │ │ │ API Key │ │ │ │ User/Team Quota │ │ │ │ RPM / TPM Limit │ │ │ │ Model ACL │ │ │ │ Routing │ │ │ └──────────┬──────────┘ │ │ │ │ │ ┌──────┴────────────────────────────────┐ │ │ │ Model Router │ │ │ │ │ │ │ │ model alias → active deployment │ │ │ └──────┬────────────────────────────────┘ │ │ │ │ │ ┌────────┴────────────────────────────────────────────────────────┐ │ │ │ GPU Serving Pool │ │ │ │ │ │ │ │ DeepSeek GLM Qwen gpt-oss PaddleOCR │ │ │ │ vLLM vLLM vLLM vLLM VLM │ │ │ │ │ │ │
velog
가능합니다. 다만 여기서는 세 층을 분리해서 설계 하는 게 좋습니다. K8s: Pod가 CPU/GPU를 같은 NUMA에 받도록 한다. NVIDIA/NCCL: 한 모델에 여러 GPU를 쓸 때 NVLink/NVSwitch를 최대한 사용한다. vLLM: TP/DP/EP와 KV cache를 workload 특성에 맞춘다. 특히 B300 8-GPU 노드라면 “무조건 single-numa-node”가 정답은 아닙니다. 8-GPU 모델은 양쪽 NUMA를 사용해야 하기 때문입니다. 1. K8s에서 CPU ↔ GPU affinity GPU 노드 kubelet의 기본 방향은 다음입니다. # /var/lib/kubelet/config.yaml cpuManagerPolicy: static topologyManagerPolicy: restricted topologyManagerScope: pod CPU Manager=static 은 Guaranteed Pod에 exclusive CPU를 할당할 수 있게 하고, Topology Manager는 CPU와 GPU 같은 device의 topology hint를 종합합니다. restricted 는 적절한 topology alignment를 만들 수 없으면 Pod admission을 거부합니다. Kubernetes 그리고 serving Pod를 반드시 Guaranteed QoS 로 만드는 게 중요합니다. resources: requests: cpu: "16" memory: "128Gi" nvidia.com/gpu: "2" limits: cpu: "16" memory: "128Gi" nvidia.com/gpu: "2" 즉 CPU의 requests == limits 로 잡습니다. 예를 들어 물리 topology가: NUMA 0 NUMA 1 ──────────────── ──────────────── CPU 0-63 CPU 64-127 GPU 0 GPU 4 GPU 1 GPU 5 GPU 2 GPU 6 GPU 3 GPU 7 NIC0 NIC1 이라면 2-GPU serving Pod에 대해 목표는: Pod A CPU 0-15 │ NUMA 0 ├── GPU0 └── GPU1 입니다. 반대로 이것은 피하고 싶은 상태입니다. CPU NUMA0 │ └──────── PCIe/NUMA crossing ─────── GPU5 NUMA1 single-numa-node 를 쓰면 더 강력하지 않나? 맞습니다. topologyManagerPolicy: single-numa-node 이면 CPU/GPU 등을 하나의 NUMA node에 배치할 수 없는 Pod는 거부합니다. Kubernetes 그런데 네 환경에서는 문제가 있습니다. 2 GPU Pod GPU0 + GPU1 → NUMA0 → OK 4 GPU Pod GPU0~3 → NUMA0 → OK 8 GPU Pod GPU0~7 → NUMA0 + NUMA1 → single NUMA 불가능 그래서 다양한 크기의 LLM serving을 같은 B300 node에서 운영한다면 restricted 를 기본으로 두는 게 더 현실적 입니다. 2. NIC locality는 Topology Manager만으로 끝나지 않는다 여기가 중요한 부분입니다. 일반적인 Kubernetes eth0 같은 host networking NIC를 사용한다고 해서 Topology Manager가 자동으로 GPU0 → NUMA0 → NIC0 GPU5 → NUMA1 → NIC1 을 만들어주는 것은 아닙니다. NIC까지 topology-managed device로 다루려면 SR-IOV/device plugin 같은 구조가 필요합니다. 따라서 네 Ethernet 구성이 예를 들어: B300 Node NUMA0 NUMA1 │ │ GPU0-3 GPU4-7 │ │ NIC0 400G NIC1 400G │ │ └────────── Ethernet Fabric ───────────┘ │ AIStor 라면 가장 강한 locality를 원할 때는: GPU Device Plugin + SR-IOV Network Device Plugin + CPU Manager + Topology Manager 식으로 CPU/GPU/NIC를 resource로 노출하는 방식을 검토합니다. 다만 AIStor에서 모델을 한 번 NVMe/HBM으로 stage한 후 serving하는 구조라면 NIC locality를 지나치게 최적화할 필요는 없습니다. 실제 inference hot path가: Client │ NIC │ CPU │ GPU │ GPU HBM 이기 때문에 request traffic 자체가 모델 weight loading traffic보다 훨씬 작을 가능성이 큽니다. NIC locality는 오히려 multi-node TP/EP/NCCL 또는 KV-cache remote offload 를 할 때 훨씬 중요해집니다. 3. GPU topology는 조금 다르게 생각해야 합니다 B300 8-GPU 시스템에서 확인해야 할 첫 번째 명령은: nvidia-smi topo -m 입니다. 목표는 대략 다음 관계를 파악하는 것입니다. NVLink / NVSwitch GPU0 ─┐ GPU1 ─┤ GPU2 ─┤ GPU3 ─┼──── NVSwitch fabric GPU4 ─┤ GPU5 ─┤ GPU6 ─┤ GPU7 ─┘ 여기서 중요한 점은 CPU NUMA와 GPU-to-GPU topology를 동일하게 생각하면 안 된다는 것 입니다. CPU에서는: CPU NUMA0 → GPU0~3 CPU NUMA1 → GPU4~7 같은 locality가 중요할 수 있지만 GPU끼리 Tensor Parallel communication을 할 때는: GPU │ NVLink │ NVSwitch │ GPU 경로가 핵심입니다. 따라서 8 GPU를 사용하는 큰 모델에서는 CPU NUMA 때문에 GPU를 억지로 4+4 별도 serving group으로 쪼개는 것보다 NVSwitch fabric 전체를 활용하는 TP/EP 구성이 더 중요할 수 있습니다. 4. 그래서 모델별 GPU 구성부터 정해야 한다 네 모델 목록을 보면 크게 Dense / MoE / Vision/OCR 로 나눠 접근하는 게 좋습니다. vLLM도 TP는 큰 모델을 여러 GPU에 shard하는 일반적인 방식이고, MoE에는 EP를 별도로 지원합니다. vLLM 예를 들어 설계 개념은: workload 우선 검토 작은 모델 GPU 1 + DP replicas 중형 Dense TP=2 또는 TP=4 큰 Dense TP=4/8 큰 MoE EP + DP/TP OCR/Vision GPU 1~2 + DP 동시 요청이 많음 DP 증가 Dense model이라면: TP=4 GPU0 ─┐ GPU1 ─┤ GPU2 ─┼── NVLink/NVSwitch GPU3 ─┘ 한 model instance 가 되고, GPU memory가 충분해서 모델 하나가 GPU 한 장에 들어가고 throughput이 중요하다면: DP=4 GPU0 → Model replica #0 GPU1 → Model replica #1 GPU2 → Model replica #2 GPU3 → Model replica #3 ▲ Load Balancer 가 더 적합합니다. vLLM의 DP는 model weight를 replica별로 가지고 독립 request batch를 처리합니다. vLLM 5. MoE는 특히 EP를 검토해야 한다 DeepSeek/Qwen 계열처럼 MoE 모델이면 이야기가 달라집니다. 예를 들어 8 GPU에서: MoE model GPU0 ─ expert set ─┐ GPU1 ─ expert set ─┤ GPU2 ─ expert set ─┤ GPU3 ─ expert set ─┼── NVSwitch GPU4 ─ expert set ─┤ GPU5 ─ expert set ─┤ GPU6 ─ expert set ─┤ GPU7 ─ expert set ─┘ 처럼 expert를 GPU에 분산시키는 Expert Parallelism 을 사용할 수 있습니다. vLLM에서는: --enable-expert-parallel 을 제공하고, 현재 EP 크기는 TP와 DP 구성으로 결정됩니다. vLLM 따라서 네 workload에서는 단순히 "8 GPU 있으니까 TP=8" 으로 통일하면 안 됩니다. Dense와 MoE의 최적 parallelism이 다릅니다. 6. KV cache가 실제 serving 성능에서 매우 중요 LLM inference에서 GPU HBM은 크게: GPU HBM ─────────────────────────── Model weights ████████████████ KV cache ████████ Activation / workspace ████ CUDA/NCCL/etc ██ 로 사용됩니다. model weight를 제외하고 남는 HBM을 얼마나 KV cache에 줄 수 있는지가 동시 request 수와 context length 에 큰 영향을 줍니다. vLLM은 이를 --gpu-memory-utilization 이나 더 직접적인 --kv-cache-memory-bytes 로 제어할 수 있습니다. 현재 기본 gpu-memory-utilization 은 0.92이고, KV cache capacity와 max concurrency도 startup log에서 확인할 수 있습니다. vLLM 처음부터 0.99 같은 공격적인 값을 쓰기보다는 실제 모델을 load해서 여유 HBM과 peak usage를 측정하는 게 좋습니다. 7. Prefix caching은 적극 검토 네가 말한 workload에는 prefix caching 도 꽤 중요할 수 있습니다. 예를 들어 OCR/document workload에서 여러 요청이: [공통 system prompt] [공통 instruction] [document A] [공통 system prompt] [공통 instruction] [document B] [공통 system prompt] [공통 instruction] [document C] 처럼 동일한 긴 prefix를 가지고 있다면: shared prefix │ KV cache █████████ / | \ / | \ doc A doc B doc C 로 reuse할 수 있습니다. vLLM에는 enable_prefix_caching 이 KV-cache configuration으로 존재합니다. vLLM 8. KV cache를 NVMe에 바로 내리는 것은 1순위가 아니다 네 시스템에는 NVMe 8개가 있지만 저는 우선: HBM │ ├── model weights └── KV cache ← 1순위 로 갑니다. HBM이 부족해졌다고 바로 HBM → NVMe 를 하기보다 먼저: HBM │ ▼ CPU RAM offload 계층을 검토하는 편이 낫습니다. 그리고 훨씬 큰 KV capacity가 필요한 경우: GPU HBM HOT │ CPU RAM WARM │ MemKV/RDMA external cache │ NVMe capacity tier 같은 계층을 검토할 수 있습니다. 현재 vLLM에는 native CPU KV offloading과 LMCache backend가 있으며 KV offload 크기를 설정할 수 있습니다. vLLM 이 부분이 앞에서 이야기했던 MemKV 테스트와 직접 연결되는 영역 입니다. 9. 네 환경에서는 전체를 이렇게 맞추는 게 좋다 제가 현재까지 나온 조건을 합치면 목표 architecture는 이쪽입니다. B300 Node ┌────────────────────────────────────────────────────┐ │ │ │ NUMA0 NUMA1 │ │ ───── ───── │ │ CPU cores CPU cores │ │ │ │ │ │ GPU0 GPU1 GPU2 GPU3 GPU4 GPU5 GPU6 GPU7 │ │ \ | | / \ | | / │ │ └──┴───┴───────────────────────┴───┴───┘ │ │ │ │ │ NVLink/NVSwitch │ │ │ │ NIC0 400G NIC1 400G │ │ │ │ │ │ └──────── Ethernet Fabric ───────┘ │ │ │ │ │ AIStor / S3 │ │ │ │ 8 x NVMe RAID0 + XFS │ │ │ │ │ /model-cache │ │ /dataset-cache │ │ /scratch │ └────────────────────────────────────────────────────┘ 그리고 Kubernetes는: CPU Manager static │ ├─────────────┐ │ │ Topology Manager │ restricted │ │ │ ▼ ▼ CPU GPU │ │ └──── NUMA ───┘ 로 시작하겠습니다. 중요한 운영 원칙 하나 GPU 1~4개짜리 serving Pod와 8-GPU serving Pod를 구분해서 생각하는 것이 좋습니다. GPU 1~4 Pod ──────────── 가능하면 CPU NUMA0 + GPU0~3 또는 CPU NUMA1 + GPU4~7 → NUMA locality 중요 GPU 8 Pod ───────── CPU NUMA0 + NUMA1 GPU0~7 NVSwitch 전체 → NUMA locality보다 NVLink/NVSwitch + TP/EP 최적화 중요 따라서 이 클러스터에서는 Topology Manager를 single-numa-node 로 고정하기보다 restricted 로 두는 것이 제가 권하는 기본안 입니다. 그리고 그 위에서 Dense 모델은 TP/DP , MoE 모델은 TP/DP/EP , KV cache는 HBM 우선 + 필요 시 CPU/MemKV offload 로 모델별 profile을 만드는 방식이 좋습니다. 네 모델 목록처럼 주간/야간 workload가 명확하다면, 다음 단계에서는 각 모델별로 GPU 수 / TP / DP / EP / 예상 HBM(weight) / KV-cache HBM / CPU core / NUMA / local NVMe cache 를 한 장짜리 배치표로 만드는 것 이 가장 유용합니다. 그 표가 나오면 “주간에는 GPU0~7을 어떻게 쪼개고, 야간에는 어떻게 재배치할지”까지 구체적으로 결정할 수 있습니다. == 맞습니다. NIC이 특정 NUMA에 명확하게 붙지 않는 topology라면 NIC까지 억지로 NUMA affinity 대상으로 묶는 것보다 CPU↔GPU locality를 우선 하는 게 좋습니다. 특히 inference 요청 트래픽 자체는 보통 model/KV 통신보다 작으므로, NIC locality는 multi-node TP/EP, remote KV cache, AIStor 대량 staging 이 있을 때 더 중요해집니다. 그리고 TP / DP / EP 는 “모델을 여러 GPU에 어떻게 나눠서 돌릴 것인가”를 정하는 서로 다른 축입니다. TP — Tensor Parallelism 한 모델의 tensor/weight 계산을 여러 GPU에 나눕니다. 예를 들어 모델 하나가 GPU 한 장에 안 들어가거나, 한 요청의 처리 속도를 높이고 싶다면: 하나의 Model │ Tensor Parallel = 4 │ ┌────────────┼────────────┐ ▼ ▼ ▼ GPU0 GPU1 ... GPU3 Weight Weight Weight 1/4 1/4 1/4 │ │ │ └───── collective ────────┘ NVLink/NVSwitch 각 transformer layer를 계산할 때 GPU끼리 자주 통신합니다. 그래서 TP는 GPU interconnect 성능에 매우 민감 합니다. B300 8-GPU NVSwitch 시스템이면 TP에 좋은 환경입니다. TP=1 → GPU간 TP communication 없음 TP=2 → GPU 2개가 하나의 model TP=4 → GPU 4개 TP=8 → GPU 8개 단순히 생각하면: TP = 모델 하나를 몇 GPU로 쪼갤 것인가 입니다. DP — Data Parallelism DP는 반대로 모델을 복제 합니다. 예를 들어 모델이 GPU 1장에 충분히 들어간다고 합시다. Load Balancer │ ┌────────────┼────────────┐ ▼ ▼ ▼ Request A Request B Request C │ │ │ GPU0 GPU1 GPU2 Model A Model A Model A replica replica replica GPU마다 동일한 model weight가 있습니다. 그래서: DP = 같은 모델을 몇 벌 만들어 요청을 병렬 처리할 것인가 입니다. Serving에서 동시 요청 처리량을 늘리는 데 매우 중요합니다. 예를 들어 8 GPU에서 모델 하나가 2 GPU를 요구한다면: TP=2 DP=4 GPU0 ─┐ ├ Model replica #1 GPU1 ─┘ GPU2 ─┐ ├ Model replica #2 GPU3 ─┘ GPU4 ─┐ ├ Model replica #3 GPU5 ─┘ GPU6 ─┐ ├ Model replica #4 GPU7 ─┘ 즉 전체 GPU 수는 간단한 경우: TP × DP 2 × 4 = 8 GPU 가 됩니다. 이 구성은 serving에서 매우 흔합니다. EP — Expert Parallelism EP는 MoE(Mixture of Experts) 모델에서만 특히 중요한 개념 입니다. MoE에서는 transformer 내부에 여러 expert가 있습니다. Token │ Router │ ┌─────────┼─────────┐ ▼ ▼ ▼ Expert 0 Expert 1 ... Expert N 모든 token이 모든 expert를 계산하는 게 아니라 router가 일부 expert를 선택합니다. 예를 들어: token A → expert 3, 17 token B → expert 8, 21 token C → expert 3, 11 처럼 됩니다. Expert가 많으면 한 GPU에 전부 넣기 어려우므로 여러 GPU에 분산합니다. EP=4 GPU0 ├ Expert 0 ├ Expert 1 └ Expert 2 GPU1 ├ Expert 3 ├ Expert 4 └ Expert 5 GPU2 ├ Expert 6 ├ Expert 7 └ Expert 8 GPU3 ├ Expert 9 ├ Expert 10 └ Expert 11 이게 Expert Parallelism 입니다. 즉: EP = MoE의 Expert들을 몇 GPU에 분산할 것인가 라고 생각하면 됩니다. TP와 EP의 결정적인 차이 둘 다 model을 여러 GPU에 나누니까 비슷해 보이지만 통신 패턴이 상당히 다릅니다. TP ──────────────────── GPU0 ←─────→ GPU1 ↑ ↑ │ AllReduce │ ↓ ↓ GPU2 ←─────→ GPU3 Tensor 계산 결과를 GPU들이 계속 교환 EP ──────────────────── Router │ token dispatch │ ┌──────────┼──────────┐ ▼ ▼ ▼ GPU0 GPU1 GPU2 Expert Expert Expert All-to-All TP는 AllReduce 계열 통신 , EP는 All-to-All token dispatch 가 중요합니다. 그래서 둘 다 NVLink/NVSwitch가 매우 유용하지만 EP는 특히 All-to-All network/interconnect 성능과 load balancing에 민감 합니다. TP + DP + EP를 동시에 사용할 수도 있습니다 예를 들어 8 GPU MoE serving을 생각해봅시다. 개념적으로: 8 GPU ┌─────────┴─────────┐ │ │ DP replica 0 DP replica 1 │ │ GPU0~3 GPU4~7 │ │ EP=4 EP=4 라면: DP = 2 EP = 4 형태입니다. 각 replica 내부의 expert를 4 GPU에 분산하고, 그런 serving replica를 두 개 만드는 겁니다. 요청은: Requests │ Load Balance / \ ▼ ▼ DP replica 0 DP replica 1 GPU0~3 GPU4~7 │ │ EP=4 EP=4 로 갑니다. 이 구조는 네 환경처럼 8 GPU B300 + NVSwitch 에서 검토할 가치가 큽니다. 그런데 TP × DP × EP = GPU 수 라고 이해하면 안 됩니다 이건 중요한 부분입니다. TP와 EP는 항상 독립적으로 GPU를 곱해서 사용하는 관계가 아닙니다. 예를 들어 vLLM에서 MoE expert parallel을 사용하면 TP
velog
🧩 결합도 · 응집도 정리 정보처리기사 실기에서 자주 출제되는 결합도(Coupling) 와 응집도(Cohesion) 를 문제 중심으로 정리한다. 📌 핵심 개념 좋은 모듈 설계의 기본은 다음과 같다. 결합도(Coupling) ↓ 낮게 응집도(Cohesion) ↑ 높게 즉, 모듈끼리는 최대한 독립적으로 모듈 내부는 하나의 목적에 집중하도록 설계하는 것이 좋다. 📌 연습 문제 1. 결합도 순서 다음 결합도 유형을 결합도가 낮은 것부터 높은 것까지 나열하시오. 공통 결합도 제어 결합도 자료 결합도 내용 결합도 외부 결합도 스탬프 결합도 풀이 결합도는 모듈끼리 얼마나 강하게 의존하는가 를 나타낸다. 결합도가 낮을수록 좋은 설계이다. 낮은 순서부터 정리하면 자료 결합도 ↓ 스탬프 결합도 ↓ 제어 결합도 ↓ 외부 결합도 ↓ 공통 결합도 ↓ 내용 결합도 ✅ 답 자료 결합도 → 스탬프 결합도 → 제어 결합도 → 외부 결합도 → 공통 결합도 → 내용 결합도 2. 자료 · 스탬프 · 제어 결합도 다음 설명에 해당하는 결합도를 순서대로 쓰시오. 1 호출할 때 고객번호와 금액처럼 필요한 값만 넘긴다. 2 주문 객체 전체를 넘기며 받은 모듈은 그중 일부 항목만 사용한다. 3 처리 모드를 나타내는 플래그를 넘겨 받은 모듈의 실행 경로를 고르게 한다. 풀이 1 필요한 값만 전달 필요한 데이터 값만 파라미터로 전달한다. 값만 전달 → 자료 결합도 2 객체나 구조체 전체 전달 필요한 값 하나가 아니라 객체, 배열, 구조체 등 자료구조 전체 를 전달한다. 자료구조 전체 전달 → 스탬프 결합도 3 제어용 플래그 전달 전달받은 값에 따라 상대 모듈의 처리 방향이 결정된다. Flag Mode 기능 선택값 → 제어 결합도 ✅ 답 1 자료 결합도 2 스탬프 결합도 3 제어 결합도 3. 외부 · 공통 · 내용 결합도 다음 설명에 해당하는 결합도를 순서대로 쓰시오. 1 여러 모듈이 외부 장치에서 정한 통신 형식에 함께 의존한다. 2 여러 모듈이 하나의 전역 설정값을 함께 읽고 수정한다. 3 한 모듈이 다른 모듈의 내부 변수와 내부 처리 위치를 직접 참조한다. 풀이 1 외부 환경에 함께 의존 외부 장치, 통신 규약, 데이터 형식 등에 여러 모듈이 함께 의존한다. 외부 장치 외부 형식 외부 프로토콜 → 외부 결합도 2 전역 변수 공유 여러 모듈이 같은 전역 데이터를 공유한다. 전역 변수 공유 → 공통 결합도 3 다른 모듈 내부를 직접 접근 다른 모듈의 내부 변수나 코드 위치를 직접 사용한다. 다른 모듈 내부 침범 → 내용 결합도 결합도가 가장 강한 형태이다. ✅ 답 1 외부 결합도 2 공통 결합도 3 내용 결합도 4. 응집도 순서 다음 응집도 유형을 응집도가 낮은 것부터 높은 것까지 나열하시오. 기능적 응집도 시간적 응집도 순차적 응집도 우연적 응집도 교환적 응집도 논리적 응집도 절차적 응집도 풀이 응집도는 하나의 모듈 안에 있는 기능들이 얼마나 밀접하게 관련되어 있는가 를 나타낸다. 응집도는 높을수록 좋은 설계 이다. 낮은 순서부터 정리하면 우연적 ↓ 논리적 ↓ 시간적 ↓ 절차적 ↓ 교환적 ↓ 순차적 ↓ 기능적 ✅ 답 우연적 응집도 → 논리적 응집도 → 시간적 응집도 → 절차적 응집도 → 교환적 응집도 → 순차적 응집도 → 기능적 응집도 5. 논리적 · 시간적 · 절차적 응집도 다음 설명에 해당하는 응집도를 순서대로 쓰시오. 1 입력 기능처럼 성격이 비슷한 작업을 한 모듈에 묶고 선택값으로 하나를 실행한다. 2 시작할 때 필요한 설정 읽기, 메모리 준비, 로그 열기를 한 모듈에 묶는다. 3 여러 작업이 정해진 순서로 실행되지만 앞 작업의 출력이 다음 작업의 입력이라는 조건은 없다. 풀이 1 비슷한 종류의 기능을 묶음 기능의 성격은 비슷하지만 실제 수행할 기능은 제어값에 따라 선택된다. 비슷한 종류 + 선택해서 실행 → 논리적 응집도 2 같은 시간대에 실행 시스템 시작 시 수행되는 작업처럼 실행 시점이 같아서 묶인다. 초기화 종료 처리 시작 시 실행 → 시간적 응집도 3 순서만 관련 있음 여러 기능이 특정 순서대로 수행된다. A → B → C 순서는 중요 데이터 전달은 필수 아님 → 절차적 응집도 ✅ 답 1 논리적 응집도 2 시간적 응집도 3 절차적 응집도 6. 교환적 · 순차적 · 기능적 응집도 다음 설명에 해당하는 응집도를 순서대로 쓰시오. 1 같은 고객 자료를 사용해 주소 표시와 등급 조회라는 서로 다른 기능을 수행한다. 2 문자열을 해석한 결과가 검증 작업의 입력으로 바로 넘어간다. 3 모듈 안의 모든 요소가 주문 금액 계산이라는 한 가지 기능만 완성한다. 풀이 1 같은 데이터 사용 서로 다른 기능이지만 동일한 입력이나 출력을 공유한다. 같은 데이터 사용 → 교환적 응집도 교환적 응집도는 통신적 응집도 라고도 한다. 2 앞의 출력 → 다음 입력 이전 단계의 처리 결과가 다음 단계의 입력으로 사용된다. A의 출력 ↓ B의 입력 → 순차적 응집도 3 하나의 기능만 수행 모듈의 모든 요소가 하나의 명확한 목적을 수행한다. 하나의 기능 → 기능적 응집도 응집도가 가장 높은 형태이다. ✅ 답 1 교환적 응집도 2 순차적 응집도 3 기능적 응집도 📚 2020~2026년 관련 기출문제 📌 2020년 1회 11번 다음은 공통 모듈 구현의 개념에 대한 설명이다. 괄호 ( ) 안에 알맞은 용어를 쓰시오. 소프트웨어 개발에 있어 기능을 분할하고 추상화하여 성능을 향상시키고 유지보수를 효과적으로 하기 위한 공통 컴포넌트 구현 기법이다. 인터페이스 모듈, 데이터베이스 접근 모듈 등 필요한 공통 모듈을 구현한다. 모듈 간의 ( 1 ) 은/는 줄이고, ( 2 ) 은/는 높은 공통 모듈 구현을 권장하고 있다. 풀이 좋은 모듈 설계는 결합도 ↓ 응집도 ↑ 가 기본 원칙이다. 따라서 모듈 간 관계 → 결합도는 낮게 모듈 내부 관계 → 응집도는 높게 ✅ 답 1 : 결합도 2 : 응집도 📌 2021년 1회 19번 다음은 결합도에 대한 설명이다. 빈칸에 들어갈 알맞은 용어를 보기에서 찾아 쓰시오. (1) 다른 모듈 내부에 있는 변수나 기능을 다른 모듈에서 사용하는 경우의 결합도 (2) 모듈 간의 인터페이스로 배열이나 객체, 구조 등이 전달되는 경우의 결합도 (3) 파라미터가 아닌 모듈 밖에 선언된 전역 변수를 참조하고 전역 변수를 갱신하는 식으로 상호작용하는 경우의 결합도 [보기] 자료 결합도 / 스탬프 결합도 / 제어 결합도 / 공통 결합도 / 내용 결합도 / 외부 결합도 풀이 1 다른 모듈 내부 접근 내부 변수 내부 기능 직접 접근 → 내용 결합도 2 배열 · 객체 · 구조체 전달 자료구조 전체 전달 → 스탬프 결합도 3 전역 변수 공유 전역 변수 → 공통 결합도 ✅ 답 1 : 내용 결합도 2 : 스탬프 결합도 3 : 공통 결합도 📌 2021년 2회 11번 응집도 문제로써, 각 번호에 해당하는 응집도를 쓰시오. 입출력 간 연관성은 없으나, 순서에 따라 수행되는 것 동일한 입력과 출력 사용 하나의 기능에 모두 기여하고 밀접하게 연관되어 있는 것 풀이 1 순서대로 수행 순서 → 절차적 응집도 2 동일한 입력과 출력 같은 데이터 사용 → 교환적 응집도 3 하나의 기능 하나의 목적 → 기능적 응집도 ✅ 답 1 : 절차적 응집도 2 : 교환적(통신적) 응집도 3 : 기능적 응집도 📌 2021년 3회 5번 다음은 Coupling에 대한 설명이다. 설명에 대한 Coupling 종류를 영문으로 작성하시오. 어떤 모듈이 다른 모듈의 내부 논리 조직을 제어하기 위한 목적으로 제어 신호를 이용하여 통신하는 경우의 결합도 하위 모듈에서 상위 모듈로 제어 신호가 이동하여 상위 모듈에게 처리 명령을 부여하는 권리 전도 현상이 발생 풀이 문제의 핵심은 제어 신호 처리 명령 내부 흐름 결정 이다. 따라서 제어 결합도 이다. ✅ 답 Control Coupling 📌 2024년 1회 3번 다음 응집도를 결합력이 강한 순서대로 나열하시오. 기능적 응집도 교환적 응집도 시간적 응집도 우연적 응집도 풀이 응집도가 높은 순서대로 보면 기능적 ↓ 교환적 ↓ 시간적 ↓ 우연적 ✅ 답 기능적 응집도 > 교환적 응집도 > 시간적 응집도 > 우연적 응집도 📌 2024년 2회 9번 다음 설명에 해당하는 응집도 종류를 쓰시오. 모듈의 출력값이 다음 기능의 입력값으로 이어지며, 앞 단계의 처리 결과가 다음 단계의 재료가 된다. 풀이 문제의 핵심은 앞 단계 출력 ↓ 다음 단계 입력 이다. 이는 순차적 응집도 이다. ✅ 답 순차적 응집도 Sequential Cohesion 📌 2024년 2회 16번 다음 설명에 해당하는 결합도를 쓰시오. 한 모듈이 다른 모듈에 제어 플래그나 기능 선택 값을 전달하고, 받은 모듈이 그 값에 따라 내부 처리 흐름을 결정한다. 풀이 Flag 기능 선택값 처리 방향 결정 이 나오면 제어 결합도 이다. ✅ 답 제어 결합도 📌 2025년 1회 12번 다음 설명에 해당하는 결합도를 1부터 3까지 순서대로 쓰시오. 1 다른 모듈 내부의 변수나 기능을 직접 사용한다. 2 모듈 사이에서 배열, 객체, 구조체 같은 자료구조 전체를 전달한다. 3 여러 모듈이 모듈 밖의 전역 변수를 함께 참조하거나 갱신한다. 풀이 내부 직접 접근 → 내용 결합도 자료구조 전달 → 스탬프 결합도 전역 변수 공유 → 공통 결합도 ✅ 답 1 내용 결합도 2 스탬프 결합도 3 공통 결합도 📌 2026년 1회 20번 다음 설명에 해당하는 응집도 유형을 순서대로 쓰시오. 1 모듈 안의 구성 요소가 정해진 순서에 따라 기능을 수행한다. 2 같은 입력과 출력을 사용해 서로 다른 기능을 수행하는 활동이 모여 있다. 3 모듈 내부의 모든 기능이 하나의 분명한 목적을 수행한다. 풀이 1 정해진 순서 순서 → 절차적 응집도 2 같은 입력과 출력 같은 데이터 → 교환적 응집도 3 하나의 목적 하나의 기능 → 기능적 응집도 ✅ 답 1 절차적 응집도 2 교환적 응집도 3 기능적 응집도 📌 2026년 2회 2번 다음 설명에 해당하는 결합도를 보기에서 골라 기호로 쓰시오. 한 모듈이 다른 모듈의 내부 데이터나 내부 제어 흐름을 직접 참조하거나 수정한다. 예를 들어 다른 모듈의 내부 코드 위치로 분기하거나, 내부에 선언된 자료를 직접 사용하는 경우이다. ᄀ. 자료 결합도 ᄂ. 스탬프 결합도 ᄃ. 제어 결합도 ᄅ. 외부 결합도 ᄆ. 내용 결합도 풀이 핵심 표현은 다른 모듈 내부 데이터 내부 제어 흐름 직접 참조 직접 수정 이다. 이는 결합도가 가장 높은 내용 결합도 이다. ✅ 답 ᄆ. 내용 결합도 📌 결합도 한눈에 보기 결합도는 낮을수록 좋다. 순서 결합도 핵심 키워드 1 자료 결합도 필요한 값만 전달 2 스탬프 결합도 배열, 객체, 구조체 전달 3 제어 결합도 Flag, 제어값 전달 4 외부 결합도 외부 장치, 형식, 프로토콜 5 공통 결합도 전역 변수 공유 6 내용 결합도 다른 모듈 내부 직접 접근 좋음 ←────────────────────→ 나쁨 자료 → 스탬프 → 제어 → 외부 → 공통 → 내용 📌 응집도 한눈에 보기 응집도는 높을수록 좋다. 순서 응집도 핵심 키워드 1 우연적 응집도 아무 관련 없이 묶음 2 논리적 응집도 비슷한 종류의 기능 3 시간적 응집도 같은 시간에 수행 4 절차적 응집도 같은 순서로 수행 5 교환적 응집도 같은 데이터 사용 6 순차적 응집도 출력 → 다음 입력 7 기능적 응집도 하나의 기능 수행 나쁨 ←────────────────────────→ 좋음 우연 → 논리 → 시간 → 절차 → 교환 → 순차 → 기능 🧠 결합도 암기 팁 결합도 순서는 자 스 제 외 공 내 로 외운다. 자 → 자료 스 → 스탬프 제 → 제어 외 → 외부 공 → 공통 내 → 내용 즉, 자료 → 스탬프 → 제어 → 외부 → 공통 → 내용 뒤로 갈수록 결합도가 강해진다. 🧠 응집도 암기 팁 응집도 순서는 우 논 시 절 교 순 기 로 외운다. 우 → 우연적 논 → 논리적 시 → 시간적 절 → 절차적 교 → 교환적 순 → 순차적 기 → 기능적 즉, 우연 → 논리 → 시간 → 절차 → 교환 → 순차 → 기능 뒤로 갈수록 응집도가 높아진다. 🔥 시험 직전 최종 암기 결합도 ↓ 낮을수록 좋음 자료 → 스탬프 → 제어 → 외부 → 공통 → 내용 응집도 ↑ 높을수록 좋음 우연 → 논리 → 시간 → 절차 → 교환 → 순차 → 기능 그리고 문제에서 키워드만 잡으면 된다. 값만 전달 → 자료 결합도 객체/구조체 전달 → 스탬프 결합도 Flag 전달 → 제어 결합도 외부 장치/형식 → 외부 결합도 전역 변수 → 공통 결합도 다른 모듈 내부 접근 → 내용 결합도 아무 관련 없음 → 우연적 비슷한 종류 → 논리적 같은 시점 → 시간적 정해진 순서 → 절차적 같은 데이터 → 교환적 출력 → 다음 입력 → 순차적 하나의 기능 → 기능적
hacker-news-frontpage
151 points · 22 comments · by simicd
Балл: 52.09Уверенность: 46%
Подробнееhacker-news-frontpage
218 points · 102 comments · by pseudolus
Балл: 52.07Уверенность: 46%
Подробнееhacker-news-frontpage
197 points · 111 comments · by jgx0
Балл: 51.84Уверенность: 46%
ПодробнееБалл: 54.38Уверенность: 49%
Балл: 54.36Уверенность: 49%
Балл: 54.36Уверенность: 49%
Балл: 54.35Уверенность: 49%