Preparing for an English proficiency exam requires more than simply studying every day. Many candidates spend weeks learning vocabulary, practising questions, watching lessons, and completing exercises, but still struggle to understand whether they are actually improving. Without a clear way to measure progress, preparation can become repetitive and unfocused. A more effective approach combines structured learning with regular assessment. Online coaching classes can provide expert guidance, while a mock test can show how well that knowledge is being applied under exam-like conditions. When these two methods are used together, candidates can identify weaknesses, measure improvement, and make better decisions about what to practise next. For students preparing for LanguageCert, this approach can be particularly useful. LanguageCert coaching provides structured instruction and feedback, while a LanguageCert mock test allows students to evaluate their performance across different exam skills. Adding a free mock test at the beginning of preparation can also establish a useful starting point. Why Tracking Progress Matters During Exam Preparation Studying for an exam without measuring progress makes it difficult to know whether your preparation strategy is working. Completing more exercises does not automatically mean your performance is improving. A student may become faster at answering questions while continuing to make the same grammar, vocabulary, comprehension, or communication mistakes. Progress tracking creates a connection between practice and performance. Instead of asking whether you have studied enough, you can ask whether your accuracy has improved, whether you are completing tasks within the required time, and whether your weak areas are becoming stronger. This approach also prevents students from spending too much time practising skills they already perform well. If Reading is consistently strong but Writing remains below the desired level, preparation time can be adjusted accordingly. The goal is not simply to complete more practice but to make measurable improvements. Start With a Baseline Mock Test The first step in tracking preparation is understanding your current level. Before beginning an intensive study plan, take a realistic free mock test or diagnostic assessment. This gives you a starting point against which future results can be compared. The baseline test should be completed under conditions that are as close as possible to the real exam. Avoid checking answers while taking the test or giving yourself unlimited time. The purpose is to discover how you perform naturally when working with exam-style questions and time pressure. The result should not be treated as a final prediction of your actual exam score. Instead, it provides useful information about your current strengths and weaknesses. A student might discover that their overall performance looks reasonable but that one particular skill is significantly weaker than the others. Use Coaching to Understand Your Results A mock test tells you what happened, but coaching can help explain why it happened. This is one of the biggest advantages of combining assessment with instruction. During online coaching classes, trainers can review common mistakes and explain the strategies needed to improve performance. Instead of simply telling a student that an answer is incorrect, effective coaching can identify the reason behind the mistake. For example, a student may lose marks because they misunderstand the question, use an inappropriate structure, struggle with vocabulary, make grammatical errors, or spend too much time on a particular task. Each problem requires a different solution. Coaching therefore turns test results into an actionable preparation plan. Students can use feedback from their assessments to decide what they should practise before taking their next test. Track Accuracy Instead of Only Overall Scores Overall scores are useful, but they do not always provide enough information about progress. A student could achieve a similar score in two different mock tests while making completely different types of mistakes. Tracking accuracy by task or skill provides a clearer picture. If a student answers 65% of questions correctly during the first assessment and 78% during a later assessment, that improvement is meaningful even if the overall score changes only slightly. The same principle applies to LanguageCert preparation. Students should monitor performance across the relevant skills and task types instead of focusing exclusively on one overall result. Over time, this creates a performance history that makes improvement easier to recognise. It also helps students identify whether their progress is consistent or whether their scores are fluctuating significantly between tests. Monitor Timing Along With Accuracy Exam preparation is not only about getting answers correct. Students must also complete tasks within the available time. Someone who achieves excellent accuracy without time pressure may still struggle when completing the actual exam. Every mock test should therefore provide an opportunity to evaluate timing. Students can record how long they spend on different sections and identify where they are losing valuable minutes. If accuracy remains high but the student repeatedly fails to complete a section, the preparation strategy needs to address speed. On the other hand, if a student finishes quickly but makes many avoidable errors, the priority should be improving accuracy. The objective is to develop a balance between speed and correctness. Regular practice makes it easier to find a working pace before exam day. How LanguageCert Coaching Supports Progress Tracking LanguageCert coaching can provide a structured framework for turning test results into measurable improvement. Rather than preparing randomly, students can follow a progression based on their current performance. A trainer can review assessment results, identify recurring weaknesses, recommend targeted exercises, and provide feedback on subsequent attempts. This creates a continuous cycle of learning, testing, reviewing, and improving. For example, a student may initially struggle with a particular language skill. After receiving coaching and completing targeted exercises, the student can take another assessment to determine whether the problem has improved. If the results show progress, preparation can move to another area. If the same weakness remains, the learning strategy can be adjusted. This makes preparation more responsive instead of following the same study routine regardless of performance. Use LanguageCert Mock Tests at Different Stages A LanguageCert mock test should not be reserved for the final days before the exam. Using mock assessments at different stages provides a much clearer view of development. An early test can establish a baseline. Tests during the middle of preparation can measure whether coaching and practice are producing improvements. Later tests can assess readiness and help students become comfortable with the complete exam experience. However, taking full-length tests every day is not necessarily the best approach. Mock tests should be combined with targeted practice. Students need enough time between assessments to analyse mistakes, study weak areas, and apply feedback. The value of a mock test comes from what happens after the test. Taking another assessment without addressing previous mistakes may simply reproduce the same problems. Create a Preparation Progress Record Keeping a simple progress record can make preparation much more organised. After every assessment, students can record their score, accuracy, timing, difficult task types, recurring mistakes, and the areas they need to practise. The purpose is not to create complicated statistics. The record should make changes in performance easy to see. For example, if a student notices that their accuracy has improved steadily but their timing remains inconsistent, the next stage of preparation can focus on timed practice. If vocabulary-related errors are decreasing but grammar mistakes remain frequent, coaching can concentrate on grammar. A progress record turns preparation into a measurable process rather than relying on memory or assumptions. Review Mistakes After Every Mock Test One of the most important stages of progress tracking happens after the test. Students should not simply look at their score and move on to another practice session. Each incorrect answer should be reviewed to determine the reason for the mistake. Some errors happen because the student does not understand the concept. Others occur because of poor concentration, rushed decisions, unfamiliar vocabulary, misunderstanding instructions, or ineffective time management. Identifying the cause is more valuable than simply counting the number of incorrect answers. A mistake that occurs repeatedly is especially important. If the same type of error appears in multiple LanguageCert mock test attempts, it indicates that additional instruction or targeted practice may be necessary. Compare Results Over Time Progress becomes clearer when results are compared over several assessments rather than judged individually. For example, a student's preparation record might show gradual improvement in accuracy, faster completion times, and fewer repeated mistakes. Even when individual scores fluctuate, the overall trend can demonstrate whether preparation is moving in the right direction. Students should avoid becoming discouraged by one lower result. Mock tests can vary in difficulty, and performance can also be affected by concentration and fatigue. Looking at multiple assessments provides a more reliable picture. The important question is whether weaknesses are gradually becoming strengths and whether performance is becoming more consistent. Combine Online Coaching Classes With Independent Practice Online coaching classes are most ef
이번 주는 malloc lab을 진행했다. 지난주 디버깅 랩과 달리, 이번에는 가이드라인이 되어 주는 책이 있어서 학습이 수월한 편이었다. 책을 읽고 코드를 구현하는 순서 로 진행하니, 무엇부터 해야 할지 몰라 막막한 순간이 줄었다. 구현 과정도 생각보다 재미있었다. 작성한 코드를 테스트하면서 확인할 수 있었고, 막히면 책을 다시 참고할 수 있었다. 지금까지 다소 방치형으로 느껴졌던 학습 방식에 비하면 꽤 친절하게 느껴졌다. 물론 수월하게 느껴졌다고 해서 구현을 모두 완료한 것은 아니다. 오히려 직접 구현할 수 있겠다는 생각이 드니 내 손으로 작성하는 부분이 늘었고, 진행 속도는 느려졌다. 아예 모르겠으면 AI에게 던져 버리기 쉬운데, 이번에는 내가 고민하고 작성하는 시간이 더 많았다. 그렇다고 AI를 사용하지 않은 것은 아니다. 디버깅이나 문제를 찾는 과정에서는 여전히 도움을 받았다. 대체로 내가 코드를 작성하고 → AI의 도움으로 문제를 확인하고 → 다시 내가 수정하는 방식 으로 진행했다. 코드를 만지면서 가장 크게 느낀 점은 각 함수가 굉장히 유기적으로 연결되어 있다는 것이었다. 기능 하나를 바꾸더라도 해당 함수만 수정하면 끝나는 게 아니라, 관련된 다른 부분까지 함께 살펴봐야 했다. 작은 변경이라고 생각했던 일이 예상보다 큰 작업으로 이어지기도 해서, 무언가를 개선하려면 생각보다 용기가 필요했다. 이번 주에는 묵시적 가용 리스트를 기반으로 First Fit과 Next Fit을 구현했다. 명시적 가용 리스트는 개념까지 공부했지만, 구현은 완료하지 못했다. 완성한 범위가 아주 넓지는 않아도, 직접 구현하고 코드를 수정해 봤다는 점에서 나름대로 만족스러운 한 주였다. 특히 코드 작성은 직접 고민하면서 해 봐야 이해도 잘되고 기억에도 오래 남는다는 것을 다시 느꼈다. 이번 주에 배운 것은 malloc의 구현 방식뿐 아니라, 시간이 조금 더 걸리더라도 직접 코드를 작성하는 과정의 가치였다.
1. 서버에 양식 샘플 업로드 코드 내 경로에 엑셀 파일 생성 DEVT30 접속 아래 코드 실행 OBJECT: STRING LVPATHSERVER,/*서버 폴더 위치*/ STRING LVPATHLOCAL;/*로컬 폴더 위치*/ LVPATHSERVER = STRSUBFIX + 'FILEUPLOAD/TEMPLATE/EDUT804NK/' + 'EDUT804NK_SAMPLE_BOOK_DOC.xlsx'; / [FILEUPLOAD/TEMPLATE/EDUT804NK/EDUT804NK_SAMPLE_BOOK_DOC.xlsx] / LVPATHLOCAL = '*' + GETUSERINFO('user.home') + GETUSERINFO('file.separator') + 'Downloads' + GETUSERINFO('file.separator') + 'EDUT804NK_SAMPLE_BOOK_DOC.xlsx'; / [*C:\Users\hmoh\Downloads\EDUT804NK_SAMPLE_BOOK_DOC.xlsx] / COPYFILE LVPATHLOCAL INTO LVPATHSERVER; : 서버 위치 지정 후, 로컬 위치 명시 그리고 `COPYFILE` 로 서버에 파일 저장 ### 2. 버튼 설정  #### 다운로드 버튼 `BTNDL.Click()` GLOBAL: STRING MSGTEXT, STRING STRSUBFIX, STRING S1, STRING S2; MESSAGE DEV C9999 WITH '양식을 다운로드 하시겠습니까?'; IF CONFIRM == 'NO' THEN RETURN; ENDIF; STRSUBFIX = SYSCOMMONNK.GETFILEPATH(SYSCOMMONNK.SESSIONINFO('COMP_CODE'), 'FILE_ROOT'); S1 = STRSUBFIX + 'EDUT804NK_SAMPLE_BOOK_DOC.xlsx'; S2 = '*' + GETUSERINFO('user.home') + GETUSERINFO('file.separator') + 'Downloads' + GETUSERINFO('file.separator') + 'EDUT804NK/EDUT804NK_SAMPLE_BOOK_DOC.xlsx'; COPYFILE S1 INTO S2; MSGTEXT = ' 엑셀 파일이 생성되었습니다.' + TOCHAR(10) + 'Path : ' + S2; MESSAGE DEV I9999 WITH MSGTEXT; 핵심코드 - `STRSUBFIX = SYS..` : 회사 코드와 FILE_ROOT 정보를 이용해서 서버에 있는 파일 저장 기본 경로를 가져옴 : `Errors: can not invoke class method:SESSIONINFO` 에러 발생, BEFORE 에 변수 선언 해줘야함 `GLOBAL: SYSCOMMONNK SYSCOMMONNK;` --- #### 업로드 버튼 1. 업로드 할 수 있는 다이얼로그 호출 CALL DIALOG EDUT804F003NK WITH LOCATION 1050,50 SIZE 400,95; SET NKRENTITEM8100 TO RESIZE; #### 파일 ᄅᄋ 업로드버튼 2. `BEFORE` (EDUT804F003NK) SELECT * FROM NKRENTITEM3100 WHERE 1 = 2 INTO TMPRENTITEM; : `NKRENTITEM8100` 하고 동일한 구조의 임시테이블 생성 3. `READEXCELFILE` (EDUT804F003NK) - `ROWCOUNT = 2;` : 헤더 제외 2번쨰 줄부터 읽기 - while로 무한반복 주다가 빈값(끝)을만나면 반복 종료 - `ACCESSXLS FILENAME READ..` 로 읽어오고, 그 값 `APPEND ROW TO TMPRENTITEM;` 로 추가하고... - 그리고 매핑 ACCESSXLS FILENAME READ SHEET 1 'A' ROWCOUNT VALUE TMPRENTITEM_BOOKID; ACCESSXLS FILENAME READ SHEET 1 'B' ROWCOUNT VALUE TMPRENTITEM_RENTAMOUNT; ACCESSXLS FILENAME READ SHEET 1 'C' ROWCOUNT VALUE TMPRENTITEM_RETURNEXPECTEDDATE; ACCESSXLS FILENAME READ SHEET 1 'D' ROWCOUNT VALUE TMPRENTITEM_RETURNDATE; ACCESSXLS FILENAME READ SHEET 1 'E' ROWCOUNT VALUE TMPRENTITEM_STATUS; ACCESSXLS FILENAME READ SHEET 1 'F' ROWCOUNT VALUE TMPRENTITEM_OVERDUEAMOUNT; ACCESSXLS FILENAME READ SHEET 1 'G' ROWCOUNT VALUE TMPRENTITEM_DAMAGEDAMOUNT; 4. OK버튼 IF TRIM(FILEPATH) == '' THEN MESSAGE DEV E9999 WITH '파일을 선택해주세요.'; RETURN; ENDIF; THIS.READEXCELFILE(SCCOMPANY,FILEPATH); CLEAR ALL NKRENTITEM8100; LOOP AT TMPRENTITEM BEGIN APPEND ROW TO NKRENTITEM8100; MOVE-CORRESPONDING TMPRENTITEM TO NKRENTITEM8100; ENDLOOP; SHUTDOWN; --- 끝 이러면 넘어옴 요약해보기 엑셀 파일 서버에 올리기 > 양식다운로드 끝 엑셀로 값 채우기 > 임시테이블 구조만 복사해서 생성 > 엑셀값 반복문으로 임시테이블 채우기 > 그 임시테이블 진짜 테이블에 집어 넣기
Advanced Multimodal 발표: 25기 김태은 구성: Recap & Bridge → MLLM Architecture(LLaVA) → Multimodal Reasoning(Multimodal-CoT) → Multimodal Generation(Latent Diffusion) → Transference(ImageBind) 지난주 멀티모달 기초가 "서로 다른 두 공간을 어떻게 붙이나"였다면, 이번 주는 붙이고 나서 무엇을 할 수 있나 였다. 발표에서 아예 "Today's Journey"로 네 단계를 먼저 깔고 시작했는데, 이 네 단계가 각각 논문 하나씩에 대응된다. 1 붙이기 (Architecture) — LLaVA. LLM에 눈을 달다 2 추론 (Reasoning) — Multimodal-CoT. 여러 단계로 생각하게 하다 3 생성 (Generation) — Latent Diffusion. 텍스트로 이미지를 만들다 4 전이 (Transference) — ImageBind. 두 모달을 넘어 여섯 모달로 그래서 이 순서 그대로 정리했다. 0. Recap — 정렬만으로는 부족하다 출발점은 지난주 선택 과제였던 Mind the Gap (arXiv:2203.02053)이다. CLIP으로 학습해도 이미지 임베딩과 텍스트 임베딩은 같은 공간 안에서 서로 다른 영역(원뿔)에 분리 되어 있다. "같은 벡터 공간"이라는 형식은 같아도 두 모달이 완전히 겹치지는 않는다. 지난주 과제에서 직접 재봤던 숫자가 그대로 근거가 된다. 사전학습 CLIP의 gap이 0.8214였고, 학습을 한 번도 하지 않은 무작위 초기화 모델이 1.1437이었다. 학습이 간격을 줄이기는 해도 없애지는 못했다. 발표는 여기서 결론을 뒤집지 않고 전제로 받아들인다. 정렬(alignment)은 출발점이다. 붙인 뒤 함께 추론하고, 생성하고, 더 많은 모달로 전이하는 것이 오늘의 주제다. gap이 남아 있다는 게 실패가 아니라는 관점이 이번 주 내용 전체를 지탱한다. 완전히 겹치지 않아도 쓸 수 있다면, 다음 질문은 "어떻게 완벽히 겹치게 하나"가 아니라 "겹치지 않은 채로 무엇을 하나"가 된다. 1. MLLM Architecture — LLM에 눈을 달다 1-1. 왜 MLLM인가 기존 멀티모달 모델의 한계를 세 줄로 정리해줬다. CLIP, ViLT, LXMERT 등은 대부분 태스크 전용 이다. VQA용, 검색용, 캡셔닝용이 따로 있다 정렬과 매칭은 잘해도 자유로운 대화나 여러 단계 추론 능력은 없다 새 태스크가 생기면 데이터를 다시 모으고 다시 학습해야 한다 반대편에는 LLM의 창발적 능력이 있다. 그래서 나오는 질문이 이거였다. "이 지능이 언어에만 갇혀 있어야 할까? 눈을 달아주면 이미지도 대화하고 추론하지 않을까?" 문제는 비용이다. LLM을 이미지까지 넣어 처음부터 다시 학습하는 건 불가능하다. 그래서 발상을 바꾼다. "잘 학습된 CLIP과 LLM을 새로 만들지 말고, 연결(connect)만 하자." 무거운 두 모델은 고정(frozen)하고 가벼운 연결자만 학습한다. 적은 비용으로 LLM의 능력을 시각으로 확장하는 구조다. 1-2. 지난주 복습 — Flamingo와 BLIP-2 사실 연결자 방식은 지난주에 이미 두 개를 배웠다. Flamingo 는 고정된 LLM 레이어 사이에 Gated Cross-Attention 을 삽입해서 LLM이 시각 정보를 참조하게 한다. BLIP-2 는 Q-Former 의 learnable query가 이미지에서 텍스트에 필요한 정보만 압축하고, 그 압축된 표현을 LLM에 전달한다. 공통점은 강력한 사전학습 Vision Encoder와 사전학습 LLM을 연결하는 것이 핵심 이라는 점이다. 그런데 발표는 여기서 한 번 더 묻는다. "연결자를 꼭 이렇게 복잡하게 만들어야 할까?" 1-3. LLaVA — 선형 층 하나 LLaVA는 Large Language and Vision Assistant의 줄임말이고, 구조가 네 조각이다. 1 Vision Encoder — CLIP ViT-L/14를 freeze 한 채로 이미지를 시각 특징 Z_v 로 인코딩한다. 이건 아직 LLM이 못 알아듣는 "비전 언어"다. 2 Projection W — 학습 가능한 선형 층 1개 . Z_v 를 LLM의 단어 임베딩 공간으로 변환한다. 3 이미지 토큰 H_v — LLM이 단어처럼 취급할 수 있는 형태가 된다. 4 Vicuna — 텍스트는 LLM 자체의 텍스트 임베딩 층을 통해 H_q 로 토큰화되고, H_v 와 concat되어 하나의 입력 시퀀스 가 된다. LLM은 앞의 모든 토큰(이미지 + 텍스트)을 참고해 답을 생성한다. 발표에서 짚어준 질문 하나가 이 구조를 이해하는 데 결정적이었다. Q. W가 이미지에만 붙는 이유는? 이미지는 LLM이 해석하지 못하는 공간에서 오니 번역이 필요하고, 텍스트는 이미 LLM의 공간에 있으니 번역이 불필요하기 때문이다. 1-4. 핵심 — 연결자는 Projection 하나뿐 Cross-Attention도, Q-Former도 아닌 단순 선형 사영 으로 시각 토큰을 LLM 입력에 나란히 넣는다. 역할은 "차원을 맞춰 같은 공간에 올려놓는 것" 이고, 발표는 이걸 기초에서 배운 Early Fusion – Type II의 실제 구현 이라고 정리했다. 지난주에 Joint Fusion Type II를 "한쪽만 가공하고 다른 쪽은 freeze"로 배웠는데, 그 추상적인 분류가 실제 코드에서 nn.Linear 하나로 나타나는 걸 보니 연결이 됐다. 실제 구현도 단순하다. 초기 LLaVA는 선형 층 1개였고, LLaVA-1.5에서 2층 MLP( Linear → GELU → Linear )로 바뀐 정도다. 1-5. 그런데 왜 이게 통했나 성능은 GPT-4 대비 상대 85.1%, ScienceQA fine-tune 시 92.53%로 당시 SOTA였다. 발표가 짚은 시사점이 더 중요하다. 강력한 사전학습 인코더와 LLM이 이미 좋은 표현을 갖고 있다 필요한 것은 둘을 잇는 정렬과 지시 데이터뿐 이다 단순함 = 확장성·재현성의 이점 → 이후 대부분의 MLLM이 이 방식으로 수렴 여기서 참고로 Chameleon (Meta, 2024)이 한 번 언급되는데, 이미지를 토큰으로 변환해 텍스트 토큰과 함께 처음부터 하나의 Transformer로 학습하는 native early-fusion 방식이다. 이해와 생성을 한 모델에서 수행한다는 점이 뒤에 다시 나온다. 2. Multimodal Reasoning — 여러 단계로 생각하게 하기 2-1. CoT를 이미지에 그냥 쓰면 Reasoning은 멀티모달의 증거로부터 지식을 여러 단계의 추론으로 조합 하는 것이다. 언어 모델에서는 Chain-of-Thought가 이미 효과적이다. "차근차근 생각해보자"를 붙이면 중간 추론 과정을 생성한 뒤 답을 낸다. 그럼 이미지가 낀 문제에서도 그대로 통할까. 안 통한다. 2-2. 문제 — Hallucinated Rationale 언어 CoT를 멀티모달에 그냥 적용하면, 특히 10억(1B) 파라미터 미만의 작은 모델은 환각된 근거를 생성한다. 발표의 자석 예시가 명확했다. vision feature 없이 돌리면 *"한 자석의 S극이 다른 자석의 S극과 가장 가깝다"* 는 잘못된 근거 를 만들어내고 오답(B)으로 간다. vision feature를 넣으면 *"한 자석의 N극이 다른 자석의 S극과 가장 가깝다"* 는 올바른 근거 가 나오고 정답(A)에 도달한다. 핵심은 모델이 답을 틀린 게 아니라 근거를 지어냈고 그 근거를 따라가다 틀렸다 는 것이다. 이미지를 제대로 반영하지 못한 추론 chain이 오답으로 이어진다. 2-3. Multimodal-CoT — 2단계로 분리 해법은 근거 생성과 정답 추론을 두 단계로 나누는 것 이다. Stage 1 — Rationale Generation 입력은 질문 + 선택지 + 이미지, 출력은 근거(rationale) 텍스트. Stage 2 — Answer Inference 입력은 질문 + 선택지 + 이미지 + 근거 , 출력은 정답. 생성된 근거를 원래 질문 뒤에 concat해서 넣는다. 핵심은 두 단계 모두에 vision feature를 주입한다는 것이다. 이미지가 있어야 근거 품질이 오르고, 좋은 근거가 정답률을 올린다. 2-4. 안에서 무슨 일이 벌어지나 발표가 수식 수준으로 네 단계를 풀어줬는데, 전부 지난주에 배운 부품들이었다. 1 Encoding — 텍스트는 Transformer 인코더로, 이미지는 ViT 인코더에 선형 사영 W_h 를 적용해 텍스트와 차원을 맞춘다. (LLaVA의 projection과 같은 역할이다) 2 Interaction — 텍스트가 이미지를 참조하는 cross-attention 이다. Q = H_language , K = V = H_vision (single-head). 해석하면 각 텍스트 토큰이 이미지의 모든 패치를 훑어보고, 자신과 관련 깊은 패치일수록 더 많이 참고해 시각 정보를 가져온다. 3 Gated Fusion — 이미지 정보를 얼마나 반영할지 게이트 λ 로 조절한다. 지난주 Gated Fusion 그대로다. 4 Decoding — H_fuse 를 Transformer decoder에 넣어 타겟(근거 또는 정답)을 생성한다. 지난주에 배운 Cross-Attention과 Gated Fusion이 이 논문의 핵심 계산 두 개였다. 기초 강의의 분류가 추상적인 정리가 아니라 실제 논문의 설계도였다는 게 이 대목에서 드러났다. 2-5. 결과 1B 미만 모델이 ScienceQA에서 당시 GPT-3.5를 약 16%p 상회 했다. 75.17% → 91.68% . 분석상으로는 환각이 줄고 수렴 속도도 빨라졌다(초반부터 더 높은 정확도). 메시지: 모델을 키우는 대신, 추론을 구조로 유도해 작은 모델로도 강한 추론을 달성한다. 관련 연구도 같이 소개됐다. POPE (arXiv:2305.10355)는 MLLM의 환각(없는 물체를 있다고 하는 것)을 Yes/No로 측정하는 벤치마크다. Visual CoT (NeurIPS 2024)는 거대 MLLM이 추론에 필요한 이미지 영역을 스스로 찾아 집중 추론하고, OpenAI o3 계열의 "Thinking with Images" 는 이미지를 crop·zoom 하는 시각적 조작 자체를 추론에 포함시킨다. 3. Multimodal Generation — 텍스트로 이미지를 만들다 3-1. Diffusion 직관 멀티모달 Generation은 CMU 분류로 번역(캡셔닝, Text→Image)·요약·창작으로 나뉘는데, 이번 주는 Text-to-Image 에 집중했다. 1 학습 (Forward + Denoise) — 이미지에 노이즈를 조금씩 더한 뒤, 노이즈를 한 스텝씩 지우는 법을 학습한다. 2 생성 (Reverse) — 순수 노이즈에서 시작해 반복적으로 지워가며 이미지를 완성한다. 한 줄로 하면, 모델 ε_θ 가 노이즈 낀 이미지 x_t 와 단계 t 를 입력받아 그때 더해진 노이즈 ε 를 맞히도록 학습한다. 노이즈를 알면 지울 수 있다. 4주차 GenAI 과제에서 F.mse_loss(predicted_noise, noise) 를 직접 짰던 게 이거였다. 3-2. Latent Diffusion — 압축과 생성을 분리한다 문제는 비용이다. 픽셀 공간에서 직접 diffusion을 돌리면 막대한 계산이 든다. 당시 선도 모델이 150~1000 V100 GPU-day 를 썼다. LDM(Rombach et al., 2022)의 핵심은 압축 단계와 생성 단계를 명시적으로 분리 한 것이다. 1 Perceptual compression — 별도로 사전 학습된 Autoencoder가 이미지 x 를 인코더 E 로 잠재 z 로 보낸다. 사람 눈에 안 보이는 고주파 디테일을 걷어내는 단계다. 2 생성 — 압축된 잠재 공간에서 Denoising U-Net이 노이즈를 반복 제거해 깨끗한 z 를 만든다. diffusion은 여기서만 일어난다. 픽셀이 아니라 z 위에서다. 이미지의 의미·구성을 학습하는 데 집중하는 semantic compression 구간이다. 3 복원 — 최종 z 를 디코더로 되돌려 이미지를 만든다. 발표가 한 줄로 요약해준 게 기억에 남는다. "크게 줄여서(1) → 그 안에서 그리고(2) → 다시 키운다(3). 그리는 내내 주문서(조건)를 본다." latent가 효율적인 이유도 두 가지로 정리됐다. autoencoder가 지각적으로 안 중요한 걸 미리 제거해서 diffusion이 의미에만 집중 할 수 있고, 한 번 학습한 범용 autoencoder를 여러 생성 모델이 재사용 할 수 있다. Stable Diffusion 생태계의 기반이 이거다. 3-3. 조건은 어떻게 넣나 — Cross-Attention "고양이 그림"이라는 텍스트 조건 y 를 어떻게 넣을까. 두 단계다. 1. τ_θ(도메인 인코더)가 필요하다. 텍스트와 이미지는 형태가 완전히 다르기 때문이다(= Heterogeneity Gap). U-Net은 글자를 못 알아듣는다. 그래서 τ_θ 가 조건 y 를 U-Net이 참조할 수 있는 벡터 τ_θ(y) 로 번역 한다. 텍스트일 때는 Transformer를 쓴다. 2. τ_θ(y) 를 U-Net에 Cross-Attention으로 삽입한다. 여기서 두 개의 "왜"가 나왔다. Why cross-attention? 길이가 다른 두 시퀀스(이미지 위치 N개 ↔ 단어 M개)를 유연하게 연결할 수 있고, 다양한 모달에 효과적이기 때문이다. 왜 딱 한 번이 아니라 매 denoise 스텝마다 참조하나? diffusion 생성은 한 번에 안 끝나고 노이즈를 여러 스텝에 걸쳐 조금씩 지운다. 조건을 처음 한 번만 주면 뒤쪽 스텝이 조건을 잊고 벗어날 수 있다. 매 스텝 참조해야 단계에 맞게 조건이 계속 반영된다. 그래서 결론이 이거였다. Cross-Attention이 멀티모달 생성 모델의 조종간이다. = 텍스트 조건이 이미지 생성을 제어하는 통로 학습 Loss는 원래 diffusion의 "노이즈 예측 오차" 식에 조건 τ_θ(y) 만 추가한 형태이고, τ_θ 와 U-Net을 하나의 loss로 함께(jointly) 학습 한다. 3-4. LDM의 범용성 조건 y 는 텍스트만이 아니다. Text → Text-to-Image ("고양이 그림") Semantic Map (영역별 라벨 지도) → 레이아웃대로 이미지 생성 Images → 저해상도에서 고해상도(super-resolution), 밑그림에서 완성 등 image-to-image Representations (클래스 라벨·특징 벡터 등) → 클래스 조건 생성 핵심은 조건을 벡터로 바꾸는 τ_θ 만 갈아끼우면, 같은 U-Net + cross-attention 구조로 이 모든 조건부 생성이 가능하다는 것이다. 3-5. (참고) 이해와 생성을 한 모델로 지금까지 이해 모델(LLaVA)과 생성 모델(LDM)은 별개였다. 최근 흐름은 하나의 모델에서 이해와 생성을 함께 하는 쪽이고, Chameleon 이 그 예다. 이미지와 텍스트를 모두 토큰으로 바꿔 하나의 Transformer로 이해와 생성을 처리한다. "MLLM이 그림도 그리게 하려면?" 이라는 질문의 최신 방향 이번 주 논문 리뷰로 고른 Show-o가 정확히 이 줄기에 있다. 4. Transference — 두 모달을 넘어 여섯 모달로 4-1. 조합 폭발 문제 CLIP은 (이미지, 텍스트) 2개 모달을 대조학습으로 정렬했다. 그럼 오디오·깊이·열화상·IMU까지 넣으려면 모든 쌍의 데이터가 필요할까? 6모달이면 15쌍이다. 4-2. ImageBind의 통찰 핵심 통찰: 이미지를 허브(hub)로 삼아, "이미지와 짝지어진 데이터"만으로 정렬한다. 논문 표현으로는 align each modality's embedding to image embeddings 다. 각 모달을 이미지(또는 비디오)와만 짝지어 학습하고, 모달끼리는 직접 짝짓지 않는다. 짝의 한쪽은 항상 이미지다. 이게 되는 이유가 데이터 쪽에 있다. 이런 쌍은 자연히 함께 존재한다(naturally paired). 동영상에는 소리와 화면이 원래 같이 있고, 1인칭 영상에는 IMU와 화면이 같이 붙어 있다. 별도 라벨링 수고 없이 수집할 수 있다. 다루는 6개 모달은 이미지, 텍스트, 오디오, 깊이(depth), 열화상(thermal), IMU이고 이미지가 중심 이다. 4-3. 어떻게 학습하고 구현하나 [학습] 각 (이미지, 모달) 쌍을 InfoNCE loss 로 정렬한다. 지난주 CLIP에서 배운 대조학습 그대로이고, 짝(positive)은 가깝게 배치 내 나머지(negative)는 멀게 민다. CLIP과의 차이는 '대상' 이다. CLIP은 텍스트 하나를 상대로 정렬했지만, ImageBind는 이미지를 여러 모달에 허브로 반복 적용 한다. [구현] 정렬이 성립하려면 모든 모달이 같은 형식의 임베딩으로 나와야 한다. 그래서 모든 모달에 Transformer(ViT류) 인코더와 모달별 projection head를 두어 공통 d차원 임베딩으로 통일 한다. 그리고 여기서 영리한 선택이 들어간다. 이미지·텍스트 인코더는 OpenCLIP으로 사전학습 후 고정(frozen) — 이미 정렬된 (이미지, 텍스트) 공간을 허브의 기준점으로 삼는다 나머지(오디오·깊이·열·IMU) 인코더만 학습 — 각자를 이 고정된 이미지 공간에 맞춰 붙인다 결과적으로 CLIP이 가진 언어-이미지 지식이 새 모달로 전이되면서 여섯 모달이 하나의 공통 공간에 모인다. 4-4. Emergent Alignment 가장 인상적이었던 부분이다. (이미지-오디오)와 (이미지-텍스트)만 학습했는데, 한 번도 같이 본 적 없는 (오디오, 텍스트)도 정렬된다. 학습하지 않은 조합의 cross-modal retrieval이 zero-shot으로 가능해지고, 서로 다른 모달의 임베딩을 더하면 그 의미들이 자연스럽게 결합된다. 응용도 바로 나온다. 임베딩 산술 조합(이미지 + 오디오로 검색)이 되고, 오디오 임베딩을 CLIP 텍스트용 DALL·E-2 디코더에 넣어 Audio-to-Image 생성 까지 된다. 개 짖는 소리를 넣으면 개 사진을 찾아온다. "이미지 인코더가 강할수록 창발 성능이 좋아진다 → 허브 품질이 전체를 좌우한다." 코드로 보면 더 간단하다. 텍스트·이미지·오디오를 각각 넣었는데 결과를 그냥 내적( @ )으로 비교 한다. 같은 공간에 있다는 증거가 그것이다. 5. 실습 — LLaVA 이미지 토큰 삽입 & ImageBind 관찰 발표 내용이 그대로 과제로 이어졌다. Part A는 LLaVA의 핵심을 손으로 재현하고, Part B는 ImageBind의 허브 아이디어를 사전학습 CLIP으로 관찰한다. 5-1. 설정 CLIP_CKPT = "openai/clip-vit-base-patch32" CLIP_HIDDEN = 768 # CLIP vision encoder가 뽑는 시각 feature 차원 (Z_v) LLM_HIDDEN = 4096 # LLM(Vicuna 등)의 단어 임베딩 차원 (H_v가 맞춰야 할 목표) 이 두 숫자가 실습 전체의 뼈대다. 768에서 4096으로 가는 것 이 과제의 전부라고 해도 된다. 5-2. 고정된 CLIP에서 패치 feature 뽑기 LLaVA가 CLIP ViT를 freeze한 채 쓰는 것을 그대로 따라간다. @torch.no_grad() def encode_image_patches(pil_img): """CLIP vision encoder로 이미지의 패치별 시각 feature를 뽑습니다. 반환: [num_patches, 768] (프로젝션 전 raw 시각 feature = Z_v)""" inputs = proc(images=pil_img, return_tensors="pt").to(device) vision_out = clip.vision_model(**inputs) # frozen ViT patches = vision_out.last_hidden_state[0, 1:, :] # CLS 제외한 패치 토큰 return patches.cpu() Z_v = encode_image_patches(sample_img) CLS 토큰을 빼고 패치 토큰만 쓰는 게 포인트다. CLIP을 검색에 쓸 때는 이미지 한 장을 벡터 하나(CLS)로 요약했는데, LLaVA는 패치 하나하나를 전부 토큰으로 넘긴다. 이미지를 "단어 하나"가 아니라 "문장"으로 취급하는 셈이다. 5-3. Projection — 선형 층 하나 class Proje