Загружаем каталог…
Загружаем каталог…
들어가며 LLM을 활용한 AI 에이전트 개발 입문」 수업 · 교재 『Do it! LLM을 활용한 AI 에이전트 개발 입문 5장~6장 실습 기록 이번 글은 5장과 6장을 묶어 정리합니다. 5장은 녹음 파일로 회의록을 만드는 AI 서기 , 6장은 이미지를 분석해 퀴즈와 듣기 문제를 만드는 AI 이미지 분석가 입니다. 5장은 소리를 글자로, 6장 마지막은 글자를 소리로 바꾸는 흐름이라 이어서 보면 재미있습니다. 일부 결과 화면은 교재 출력을 바탕으로 재구성했습니다. 전체 진행 순서 5장 AI 서기 openai_api_whisper.ipynb : 위스퍼 API로 음성 → 텍스트 huggingface_whisper.ipynb : 위스퍼 모델을 로컬에서 실행, 시간대별 문장을 CSV로 저장 speaker_diarization.ipynb : pyannote로 화자 분리 whisper_stt.py : 받아쓰기와 화자 분리를 합치는 파일 summarize_and_correct.ipynb : GPT로 요약·교정, 마크다운·워드 출력 6장 AI 이미지 분석가 image_explanation.ipynb : GPT 비전으로 이미지 설명, 한계 확인 image_quiz.py : 이미지 퀴즈와 영어 듣기 문제 출제 tts.ipynb : 영어 문제를 TTS로 MP3 만들기 project_agent_20260305/ ├── chap05/ │ ├── audio/ │ │ ├── lsy_audio_2023_58s.mp3 ← 1분짜리 강의 녹음 │ │ └── 싼기타_비싼기타.mp3 ← 두 사람의 토론 녹음 │ ├── openai_api_whisper.ipynb │ ├── huggingface_whisper.ipynb │ ├── speaker_diarization.ipynb │ ├── whisper_stt.py │ └── sec04/summarize_and_correct.ipynb ├── chap06/ │ ├── data/ ← 분석할 이미지, 퀴즈 결과 파일 │ ├── image_explanation.ipynb │ ├── image_quiz.py │ └── tts.ipynb └── venv/ 1. 위스퍼 API로 음성을 텍스트로 — openai_api_whisper.ipynb 실습 음성( lsy_audio_2023_58s.mp3 )은 교재 깃허브에서 받아 chap05/audio 에 넣었습니다. 이번 장부터는 셀 단위로 실행하고 결과를 바로 볼 수 있는 주피터 노트북 을 씁니다. VS Code에서 .ipynb 파일을 만들고 Select Kernel 에서 venv 를 고르면 되고, 처음 실행할 때 나오는 ipykernel 설치 안내는 [Install]을 누르면 됩니다. 클라이언트 준비 from openai import OpenAI from dotenv import load_dotenv import os load_dotenv() api_key = os.getenv('OPENAI_API_KEY') client = OpenAI(api_key=api_key) .env 의 API 키로 OpenAI 클라이언트를 만듭니다. 노트북은 실행한 셀의 변수가 남아 있어서, 한 번만 실행해 두면 아래 셀에서 client 를 계속 쓸 수 있습니다. MP3를 텍스트로 audio_file_path = './audio/lsy_audio_2023_58s.mp3' # MP3 파일 경로 입력 with open(audio_file_path, 'rb') as audio_file: transcription = client.audio.transcriptions.create( model="whisper-1", file=audio_file ) transcription 음성 파일은 바이너리라 'rb' 모드로 열고, client.audio.transcriptions.create() 에 whisper-1 모델과 파일을 넘기면 받아쓰기 결과가 옵니다. 결과가 한 줄로 길게 나와서 textwrap 으로 60자씩 줄바꿈해 출력하도록 바꿨습니다. from openai import OpenAI from dotenv import load_dotenv import os import textwrap load_dotenv() api_key = os.getenv("OPENAI_API_KEY") client = OpenAI(api_key=api_key) audio_file_path = './audio/lsy_audio_2023_58s.mp3' with open(audio_file_path, 'rb') as audio_file: transcript = client.audio.transcriptions.create( file=audio_file, model="whisper-1" ) # 결과 출력 # print(transcript.text) # 결과 출력 (문단 형태) text = transcript.text formatted_text = "\n\n".join( textwrap.fill(paragraph, width=60) for paragraph in text.split("\n\n") ) print(formatted_text) transcript.text 로 글자만 꺼내고, textwrap.fill(width=60) 으로 줄을 맞춥니다. 1분짜리 음성이 몇 초 만에 거의 정확한 문장으로 나왔습니다. 번역하며 받아쓰기 import textwrap # Open audio file and request transcription (translation to English) with open(audio_file_path, 'rb') as audio_file: transcription = client.audio.translations.create( file=audio_file, model="whisper-1" ) # Extract text from transcription result text = transcription.text # (Fixed: transcript → transcription) # Format text into paragraphs with max 60 characters per line formatted_text = "\n\n".join( textwrap.fill(paragraph, width=60) for paragraph in text.split("\n\n") ) # Print formatted result print(formatted_text) transcriptions 를 translations 로 바꾸면 한국어 음성이 영어 텍스트로 나옵니다. 주석의 (Fixed: ...) 는 변수 이름을 잘못 써서 NameError 가 났던 걸 고친 흔적입니다. 참고 translations 는 영어로 번역하는 것만 지원합니다. 2. 위스퍼를 로컬에서 실행 — huggingface_whisper.ipynb API는 쓸 때마다 요금이 나가서, 이번에는 허깅페이스에서 openai/whisper-large-v3-turbo 모델을 내려받아 내 컴퓨터에서 돌려 봅니다. 패키지 설치 %pip install --upgrade pip %pip install --upgrade transformers datasets[audio] accelerate %pip 은 노트북의 현재 커널(venv)에 설치하는 명령입니다. 모델 실행용 transformers , 오디오 처리용 datasets[audio] , 메모리 최적화용 accelerate 를 설치합니다. 이 상태로 예제를 실행하면 FFMPEG가 없다는 오류가 납니다. FFMPEG 경로 등록 https://www.gyan.dev/ffmpeg/builds 에서 ffmpeg-git-full.7z 를 받아 압축을 풀고, bin 폴더를 PATH에 추가합니다. import os os.environ["PATH"] += os.pathsep + r"C:\github\gpt_agent_2025_book\ffmpeg-2025-01-22-full_build\bin" 지금 커널에서만 유효해서 커널을 재시작하면 다시 실행해야 합니다. 경로는 본인이 압축을 푼 위치로 바꿔 주세요. 모델 실행 import torch from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline # from datasets import load_dataset device = "cuda:0" if torch.cuda.is_available() else "cpu" torch_dtype = torch.float16 if torch.cuda.is_available() else torch.float32 model_id = "openai/whisper-large-v3-turbo" model = AutoModelForSpeechSeq2Seq.from_pretrained( model_id, torch_dtype=torch_dtype, low_cpu_mem_usage=True, use_safetensors=True ) model.to(device) processor = AutoProcessor.from_pretrained(model_id) pipe = pipeline( "automatic-speech-recognition", model=model, tokenizer=processor.tokenizer, feature_extractor=processor.feature_extractor, torch_dtype=torch_dtype, device=device, return_timestamps=True, chunk_length_s=10, stride_length_s=2, ) # dataset = load_dataset("distil-whisper/librispeech_long", "clean", split="validation") # sample = dataset[0]["audio"] sample = "./audio/lsy_audio_2023_58s.mp3" result = pipe(sample) # print(result["text"]) print(result) 허깅페이스 예제에서 바꾼 부분은 파이프라인 옵션 세 개입니다. return_timestamps=True : 문장마다 시작·끝 시간을 함께 받습니다. 뒤에서 화자 분리와 맞출 때 필요합니다. chunk_length_s=10 : 음성을 10초씩 잘라 처리합니다. stride_length_s=2 : 조각이 2초씩 겹치게 해서 경계에서 단어가 잘리는 걸 줄입니다. GPU가 있으면 cuda , 없으면 cpu 로 돌아갑니다. 실행하니 이번에는 No module named 'torch' 오류가 나서, pytorch.org에서 안내하는 명령으로 파이토치를 설치했습니다. (venv) > pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121 GPU가 없다면 사이트에서 CPU 를 골라 나오는 명령을 쓰면 됩니다. 설치 후 실행하면 전체 텍스트와 시간대별 조각( chunks )이 함께 나옵니다. chunks를 CSV로 저장 # chunks를 CSV 파일로 저장 start_end_text = [] for chunk in result["chunks"]: start = chunk["timestamp"][0] end = chunk["timestamp"][1] text = chunk["text"] start_end_text.append([start, end, text]) import pandas as pd df = pd.DataFrame(start_end_text, columns=["start", "end", "text"]) df.to_csv("lsy_audio_2023_58.csv", index=False, sep="|") display(df) 조각마다 [시작, 끝, 문장] 을 모아 데이터프레임으로 만들고 저장합니다. 문장에 쉼표가 들어갈 수 있어서 구분자는 | 를 썼습니다. 3. 화자 분리 — speaker_diarization.ipynb 받아쓰기만으로는 누가 말했는지 알 수 없어서, pyannote/speaker-diarization-3.1 모델로 화자를 나눕니다. 약관 동의가 필요한 모델이라 허깅페이스에서 액세스 토큰 을 만들고 모델 페이지에서 동의해야 합니다. 설치 %pip install pyannote.audio %pip install numpy==1.26 pyannote가 numpy 2.x와 맞지 않아서 numpy==1.26 으로 고정했습니다. 파이프라인 불러오기 # instantiate the pipeline from pyannote.audio import Pipeline pipeline = Pipeline.from_pretrained( "pyannote/speaker-diarization-3.1", use_auth_token="HUGGINGFACE_ACCESS_TOKEN_GOES_HERE" ) # CUDA를 사용할 수 있다면 CUDA를 사용하도록 설정 if torch.cuda.is_available(): pipeline.to(torch.device("cuda")) print('cuda is available') else: print('cuda is not available') use_auth_token 에 발급한 토큰을 넣습니다. 이 셀은 torch 를 쓰는데 import가 없으니, 새로 시작했다면 import torch 를 먼저 해 주세요. 실행하니 아래 오류가 났습니다. Could not download 'pyannote/segmentation-3.0' model. It might be because the model is private or gated so make sure to authenticate. ... AttributeError: 'NoneType' object has no attribute 'eval' 3.1 모델이 내부에서 쓰는 pyannote/segmentation-3.0 모델에도 따로 동의해야 했습니다. 해당 페이지에서 [Agree and access repository]를 누르면 해결됩니다. 주의 토큰은 비밀번호와 같아서, 깃허브에 올릴 때는 .env 로 빼 두는 게 안전합니다. 화자 분리 실행 # run the pipeline on an audio file diarization = pipeline("./audio/싼기타_비싼기타.mp3") # dump the diarization output to disk using RTTM format with open("./audio/싼기타_비싼기타.rttm", "w", encoding='utf-8') as rttm: diarization.write_rttm(rttm) 토론 음성 싼기타_비싼기타.mp3 로 실행하고 결과를 RTTM 파일로 저장합니다. 한 줄은 "몇 초부터 몇 초 동안 누가 말했는지"입니다. 네 번째 값이 시작 시간, 다섯 번째 값이 길이 입니다. 판다스로 정리 import pandas as pd rttm_path = "./audio/싼기타_비싼기타.rttm" df_rttm = pd.read_csv( rttm_path, # rttm 파일 경로 sep=' ', # 구분자는 띄어쓰기 header=None, # 헤더는 없음 names=['type', 'file', 'chnl', 'start', 'duration', 'C1', 'C2', 'speaker_id', 'C3', 'C4'] ) display(df_rttm) 띄어쓰기로 구분된 RTTM을 읽고, 열 이름을 직접 붙였습니다. # start + duration을 end로 변환 df_rttm['end'] = df_rttm['start'] + df_rttm['duration'] display(df_rttm) 끝 시간이 없으니 시작 + 길이로 만듭니다. df_rttm["number"] = None # number 열 만들고 None으로 초기화 df_rttm.at[0, "number"] = 0 display(df_rttm) for i in range(1, len(df_rttm)): if df_rttm.at[i, "speaker_id"] != df_rttm.at[i-1, "speaker_id"]: df_rttm.at[i, "number"] = df_rttm.at[i-1, "number"] + 1 else: df_rttm.at[i, "number"] = df_rttm.at[i-1, "number"] display(df_rttm.head(10)) 같은 사람이 이어서 말한 줄은 같은 번호, 화자가 바뀌면 번호를 1 올립니다. 숨 쉬는 틈마다 쪼개진 줄을 발언 차례 단위로 묶기 위한 준비입니다. df_rttm_grouped = df_rttm.groupby("number").agg( start=pd.NamedAgg(column='start', aggfunc='min'), end=pd.NamedAgg(column='end', aggfunc='max'), speaker_id=pd.NamedAgg(column='speaker_id', aggfunc='first') ) display(df_rttm_grouped) 번호별로 묶어서 시작은 최솟값, 끝은 최댓값, 화자는 첫 값을 씁니다. df_rttm_grouped["duration"] = df_rttm_grouped["end"] - df_rttm_grouped["start"] df_rttm_grouped = df_rttm_grouped.reset_index(drop=True) display(df_rttm_grouped) df_rttm_grouped.to_csv( "./audio/싼기타_비싼기타_rttm.csv", sep=',', index=False ) 발화 시간을 다시 계산하고 인덱스를 정리해 CSV로 저장합니다. 4. 받아쓰기 + 화자 분리 합치기 — whisper_stt.py 노트북에서 확인한 기능을 함수로 묶은 최종 파일입니다. 전체 코드 import os import torch import pandas as pd from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline from pyannote.audio import Pipeline os.environ["PATH"] += os.pathsep + r"C:\github\gpt_agent_2025_book\ffmpeg-2025-01-22-full_build\bin" # 자신이 설치한 위치로 경로 수정 def whisper_stt( audio_file_path: str, output_file_path: str = "./output.csv" ): device = "cuda:0" if torch.cuda.is_available() else "cpu" torch_dtype = torch.float16 if torch.cuda.is_available() else torch.float32 model_id = "openai/whisper-large-v3-turbo" model = AutoModelForSpeechSeq2Seq.from_pretrained( model_id, torch_dtype=torch_dtype, low_cpu_mem_usage=True, use_safetensors=True ) model.to(device) processor = AutoProcessor.from_pretrained(model_id) pipe = pipeline( "automatic-speech-recognition", model=model, tokenizer=processor.tokenizer, feature_extractor=processor.feature_extractor, torch_dtype=torch_dtype, device=device, return_timestamps=True, # 청크별로 타임스탬프를 반환 chunk_length_s=10, # 입력 오디오를 10초씩 나누기 stride_length_s=2, # 청크가 2초씩 겹치도록 나누기 ) result = pipe(audio_file_path) df = whisper_to_dataframe(result, output_file_path) return result, df def whisper_to_dataframe(result, output_file_path): s
То, что RADAR обнаружил и классифицировал для этой возможности. Это опубликованный источником текст, а не подтверждение, что предложение ещё действует.
[시스템 프로그래밍] 위스퍼·화자 분리·GPT 비전·TTS로 만드는 AI 서기와 AI 이미지 분석가. 들어가며 LLM을 활용한 AI 에이전트 개발 입문」 수업 · 교재 『Do it! LLM을 활용한 AI 에이전트 개발 입문 5장~6장 실습 기록 이번 글은 5장과 6장을 묶어 정리합니다. 5장은 녹음 파일로 회의록을 만드는 AI 서기 , 6장은 이미지를 분석해 퀴즈와 듣기 문제를 만드는 AI 이미지 분석가 입니다. 5장은 소리를 글자로, 6장 마지막은 글자를 소리로 바꾸는 흐름이라 이어서 보면 재미있습니다. 일부 결과 화면은 교재 출력을 바탕으로 재구성했습니다. 전체 진행 순서 5장 AI 서기 openai_api_whisper.ipynb : 위스퍼 API로 음성 → 텍스트…
Открыть источник