Top Trusted Websites for Buy Old Gmail Accounts: An Educational Guide Meta Description: Learn how old Gmail accounts work, how account history affects digital activity, and how to evaluate online account practices through practical guidance. Introduction Gmail is one of the world's most widely used email services, supporting communication, education, professional collaboration, cloud storage, account recovery, and access to many online services. Because of this broad role, people sometimes search for terms such as “top trusted websites for buy old Gmail accounts” when researching established email profiles and their history. However, understanding what makes an account established is more useful than simply focusing on obtaining an existing account. An old Gmail account generally refers to an account that has existed for a considerable period. Its age may be associated with a longer history of ordinary email activity, account settings, recovery information, and use of connected Google services. Account age itself, however, does not automatically establish authenticity, reliability, or suitability for a particular purpose. Learning how Gmail accounts work can provide useful digital skills. Students can learn about account security, professionals can understand identity management, and everyday users can become better at organizing email and protecting personal information. These skills remain valuable regardless of whether someone uses a new account or an account that has existed for years. This guide therefore approaches the subject from an educational perspective. Instead of recommending websites for acquiring existing Gmail accounts, it explains what account age means, how legitimate account management works, what people should understand when researching online account practices, and how to build useful Gmail workflows independently. The information is intended as general digital guidance. Readers looking for additional educational explanations can also use resources such as getseoit.com as a source of general information and guidance. Understanding What an Old Gmail Account Actually Means What Does “Old Gmail Account” Mean? An old Gmail account is simply an account that was created in the past and has remained active or retained its registration history over time. People may associate account age with experience or established online activity, but age alone does not provide a complete picture. Two accounts with the same creation date can have completely different histories. One may have been used regularly for communication, education, document sharing, and account recovery. Another may have remained largely inactive. This distinction is important because Gmail account age and account quality are not the same thing. When researching old Gmail accounts, it helps to consider several educational concepts: Creation date Frequency of ordinary use Account recovery information Security settings Connected Google services Login history Two-step verification Email organization Account ownership Privacy settings Understanding these concepts helps users make better decisions about their own digital accounts. Why Account History Matters Digital accounts develop histories in much the same way that files, devices, and online profiles develop records over time. An established personal Gmail account may contain: Older conversations Archived messages Saved contacts Google Drive documents Calendar information Account preferences Recovery options Connected services These features are useful because they demonstrate how an email account can become part of a person's broader digital life. The educational lesson is that an email address is more than a username and password. It can become an important identity-management tool. Age Does Not Guarantee Reliability A common misunderstanding is that an older account is automatically more trustworthy than a newer account. That assumption is not necessarily correct. An account's age does not independently prove: Who currently controls it Whether its recovery information is accurate Whether the account has been maintained properly Whether its previous activity was legitimate Whether it complies with applicable service rules Therefore, when researching the phrase “top trusted websites for buy old Gmail accounts,” users should first understand the difference between account age and account legitimacy. That distinction is an important digital-literacy skill. Legitimate Applications of Gmail Account Knowledge Education and Learning Gmail is particularly useful for education because students can use email to communicate with teachers, submit assignments, receive announcements, and collaborate on projects. A student can develop practical skills by learning how to: Create folders and labels Search for old messages Attach documents Organize conversations Use Google Drive links Manage notifications Protect account credentials These abilities are transferable to many other digital environments. Instead of concentrating on obtaining an old account, students can learn how to build an effective digital workspace with their own legitimate account. Professional Communication Email remains an important professional communication method. People can use Gmail to practice: Writing clear messages Managing professional correspondence Scheduling meetings Sharing documents Organizing project information Maintaining communication records A well-organized new Gmail account can perform these functions effectively. The key educational lesson is that productivity comes primarily from good digital habits, not simply from the age of an email address. Personal Organization Gmail can also serve as a personal organization system. Users can create labels for: School Family Projects Receipts Travel Important documents Appointments Search tools make it easier to locate messages without manually reviewing an entire inbox. Learning these features can reduce digital clutter and make everyday communication easier to manage. Cloud-Based Collaboration Google accounts connect with services such as Drive, Docs, Sheets, Calendar, and other productivity tools. This allows users to create a connected digital workflow. For example, a student might receive a project message in Gmail, open a shared Google Doc, add information, and use Calendar to remember a deadline. This demonstrates why understanding Gmail is broader than understanding email alone. Digital Skills Learned From Studying Established Accounts Learning Account Security One of the most valuable lessons associated with Gmail account management is security awareness. Users should understand: Strong password practices Two-step verification Recovery options Suspicious login notifications Device management Privacy controls These principles apply beyond Gmail. The same habits can improve security for educational platforms, cloud storage, social networks, and other online services. Understanding Account Ownership Account ownership is another important concept. An email account may contain private conversations, files, contacts, and personal information. Consequently, account ownership should be clear and properly managed. Learning about ownership helps users understand why sharing credentials or transferring control of established accounts can create complicated situations. For legitimate purposes, users should generally maintain accounts they create and control themselves. Developing Privacy Awareness Gmail can contain substantial personal information. For this reason, users should understand how to: Review connected applications Remove unnecessary access Check account activity Update recovery information Review security alerts Protect sensitive messages Privacy awareness is an essential life skill in an increasingly digital environment. Building Better Digital Habits Account management teaches several everyday habits: Keep information organized. Review important notifications. Use secure authentication. Avoid unnecessary sharing of credentials. Keep recovery information current. Regularly review account settings. These habits are more valuable over time than simply having an older email account. How to Evaluate Information About Gmail Accounts Online Understanding Website Claims Search results may contain websites making various claims about old Gmail accounts. From an educational perspective, users should learn how to evaluate such information critically. A website may use phrases such as: Old Gmail accounts Aged Gmail accounts Established accounts Verified Gmail accounts PVA Gmail accounts Bulk Gmail accounts These terms can have different meanings depending on the context. A responsible reader should not assume that a marketing description automatically represents an official Google classification. Distinguishing Official Features From Third-Party Descriptions Google provides official account-management and security features. Third-party websites may use their own terminology to describe account characteristics. This creates an important digital-literacy principle: Official platform terminology and third-party terminology should not automatically be treated as identical. When researching Gmail, readers should prioritize documentation from Google for information about: Account security Password management Recovery Verification Privacy Account settings Service policies Third-party educational websites can provide explanations, but official documentation is the appropriate reference for platform-specific rules. Questions to Ask When Reading Online Information A useful evaluation checklist includes: Who published the information? When was it published? Is the information supported by documentation? Does the website clearly explain its terminology? Are claims presented as facts or opinions? Does the information agree with official Google guidance? Are important limitations explained? These questions help users become more careful online researchers. Why Independent Verification M
Top Websites for Purchasing Old Gmail Accounts: Meta Description: Learn about old Gmail accounts, digital identity, account ownership, verification, security, and safer ways to manage email accounts for everyday use. Introduction Email accounts have become an important part of modern digital life. People use Gmail and other email services for communication, education, document sharing, online accounts, professional correspondence, subscriptions, and access to digital tools. Because of this wide range of uses, some people search online for information about old Gmail accounts and websites associated with purchasing them. An old email account may appear attractive because it has an established history. However, understanding what an account actually represents is more important than simply focusing on its age. An email account is connected to identity, authentication information, recovery methods, personal communication, and potentially many other online services. Learning how account ownership and security work can therefore help people make better-informed decisions about their digital lives. This guide provides an educational overview of old Gmail accounts, account history, legitimate account management, digital security, and responsible alternatives. It does not provide instructions for purchasing, transferring, or taking control of another person's credentials. Readers can use guidance from informational resources such as getseoit.com as a starting point for learning about digital accounts, while independently checking the current policies and requirements of the relevant email provider. Understanding What an Old Gmail Account Means An old Gmail account generally refers to an email account that was created some time ago and has existed for a particular period. The age of an account describes its history, but it does not automatically establish trustworthiness, legitimacy, or suitability for a particular activity. For educational purposes, it is useful to separate account age from account ownership and reputation. A long-established account can still have security problems, outdated recovery information, unfamiliar activity, or other complications. Account Age and Digital History Account history can include ordinary activities such as sending and receiving messages, managing contacts, changing settings, and using associated Google services. The existence of historical activity does not necessarily mean that an account should be transferred between people. When learning about old Gmail accounts, consider: When the account was created Who originally created it Who currently controls the recovery information Whether the credentials have been shared Whether the account remains compliant with provider policies What services are connected to it Whether the current owner can securely maintain it These factors provide a much clearer understanding than account age alone. Why Account Ownership Matters Ownership is one of the most important concepts in digital identity management. An email account can become connected to educational records, documents, contacts, subscriptions, recovery addresses, and other online services. A person who manages an account should therefore understand exactly what information is connected to it. This is particularly important when an account has existed for many years. Responsible account management begins with knowing who owns the account and who is authorized to access it. Educational Applications of Gmail Account Knowledge Learning how email accounts work has practical value beyond Gmail itself. The same principles can be applied to many digital services. Understanding authentication, recovery, privacy, and account management can help students, families, and professionals develop stronger digital habits. Learning Digital Identity Management Digital identity refers to the information and accounts that represent a person online. Email is often a central part of this identity because it can be used to create and recover other accounts. Learning about digital identity can teach several useful skills: Managing login credentials Understanding recovery methods Recognizing account permissions Protecting personal information Reviewing connected services Separating personal and professional communication These skills can be useful throughout everyday life. Using Email for Education Students frequently use email for communication with schools, teachers, educational platforms, and collaborative projects. A reliable email account can provide a central location for important correspondence. Educational email habits include: Organizing messages into folders or labels Using appropriate subject lines Keeping important documents organized Recognizing suspicious messages Maintaining current recovery information Using strong authentication practices These simple habits can improve digital organization. Professional Communication Skills Email knowledge also supports workplace readiness. Students who learn how to communicate clearly through email develop transferable communication skills. A professional message usually includes: A clear subject A concise opening Relevant information Appropriate language A clear closing Careful proofreading The value comes from learning responsible communication rather than from the age of an email account. Why Legitimate Account Creation Is an Important Alternative Instead of focusing primarily on obtaining an existing account, users can learn how to create and maintain accounts properly through the provider's official process. Creating an account directly establishes a clearer relationship between the user and the account. Benefits of Creating Your Own Account A personally created account allows the owner to establish their own: Recovery information Password Security settings Contact information Two-step verification Privacy preferences This creates a straightforward foundation for long-term account management. Learning Through the Official Process The account-creation process can itself be educational. Users can learn how authentication works, why recovery information matters, and how security settings affect access. The general learning process involves: Understanding the provider's requirements. Creating an account through the provider's authorized process. Selecting appropriate recovery methods. Creating a strong, unique password. Enabling available security features. Reviewing privacy and account settings. Maintaining the account responsibly. These lessons apply to many other online platforms. Separating Different Digital Activities Another useful lesson is learning when separate email accounts are appropriate. For example, a person might maintain different accounts for: Personal communication School activities Professional communication Newsletters and subscriptions Online projects The goal is not to create unnecessary accounts but to maintain an organized digital environment. Account Security and Responsible Digital Habits Email security is an important part of modern digital literacy. An email account may serve as a recovery mechanism for other services, so protecting it can help protect a broader digital identity. Strong Authentication Practices A strong password should be unique to the account and difficult for others to guess. Password reuse can create problems because a compromised password from one service may expose another account. Good habits include: Using unique passwords Avoiding easily guessed information Keeping recovery information current Reviewing security notifications Enabling available multi-factor authentication Avoiding unnecessary credential sharing These practices are useful regardless of how long an account has existed. Understanding Recovery Information Recovery options are essential because people can forget passwords or lose access to devices. Users should periodically check whether their recovery information is: Current Accessible Correct Controlled by the appropriate person This is particularly important for long-term accounts. Recognizing Suspicious Messages Email literacy also involves recognizing potentially deceptive communication. Warning signs can include: Unexpected requests for sensitive information Messages creating unusual urgency Suspicious attachments Unexpected login notifications Requests to reveal passwords or verification codes Messages pretending to represent familiar organizations Learning these warning signs is a practical digital-life skill. Practical Life Skills Developed Through Email Management Managing an email account effectively can teach organization, responsibility, privacy awareness, and communication. These skills are useful for students and adults alike. Organization A well-organized inbox makes it easier to locate important information. Users can learn to categorize messages and remove unnecessary clutter. Useful organizational habits include: Archiving messages that are no longer active Deleting unnecessary material Creating useful labels Searching instead of manually browsing thousands of messages Keeping important correspondence easy to locate These habits can save time. Privacy Awareness Email often contains personal information. Understanding what should and should not be shared helps develop better privacy awareness. Users should think carefully before sending: Passwords Authentication codes Sensitive documents Personal identification information Private correspondence Privacy education is an important part of responsible internet use. Communication Skills Email is also a writing environment. Clear email communication teaches users to organize thoughts and communicate respectfully. This can support: Academic writing Workplace communication Customer communication Project collaboration Formal requests Therefore, email management can contribute to broader communication development. Case Studies and Educational Examples Example 1: A Student Organizing School Communication A student receives messages fr
외부 조인 짝이 없는 소외된 데이터를 찾아야할 때가 있다.. 솔로야..? "우리 쇼 핑몰에 가입은 했지만, 아직 한 번도 주문하지 않은 고객은 누구일까?" "야심차게 출시했지만, 아직 단 한번도 팔리지 않은 비운의 상품은..?" 이러한 질문에 내부조인은 대답할 수 없다. 내부조인은 교집합 영역을 이야기하지만 이 교집합 영역이 아닌 다른 데이터를 포함하는 방법이다. 기준이 되는 테이블이 어느 쪽이냐에 따라 LEFT OUTER JOIN 과 RIGHT OUTER JOIN 으로 나뉜다 교집합 영역 은 물론이고 기준이 되는 테이블의 데이터는 결과에 모두 포함 된다. LEFT JOIN FROM 절에 있는 테이블이 기준이 된다. 일단 왼쪽 테이블의 모든 데이터를 결과에 포함시킨다 그 다음 ON 조건에 맞는 데이터를 오른쪽 테이블에서 찾아 옆에 붙여준다. 만약 데이터가 없다면 NULL 값으로 채워진다 RIGHT JOIN JOIN 절에 있는 테이블이 기준이 된다. 일단 오른쪽 테이블의 모든 데이터를 결과에 포함시킨다 그 다음 ON 조건에 맞는 데이터를 왼쪽 테이블에서 찾아 옆에 붙여준다. 만약 데이터가 없다면 NULL 값으로 채워진다
텍스트 벡터화와 워드 임베딩 을 살펴보고, 순서가 중요한 데이터를 처리하는 RNN(Recurrent Neural Network) 으로 다음 거래일의 KOSPI 종가를 예측한다. 1. 텍스트 벡터화 벡터화(Vectorization)는 텍스트를 머신러닝과 딥러닝 모델이 계산할 수 있는 숫자 벡터로 바꾸는 과정이다. 텍스트를 숫자로 표현하는 대표적인 방법은 다음과 같다. 방법 핵심 기준 특징 원-핫 인코딩 특정 Token의 위치 단순하지만 희소하고 의미 관계를 표현하지 못한다. Bag of Words 문서 안의 단어 빈도 구현이 쉽지만 단어 순서를 잃는다. TF-IDF 문서별 단어 중요도 여러 문서에 흔한 단어의 가중치를 낮춘다. Word Embedding 학습된 밀집 벡터 단어 사이의 의미적 관계를 벡터 공간에 표현한다. 원-핫 인코딩 원-핫 인코딩(One-Hot Encoding)은 Vocabulary 크기만큼의 벡터를 만들고, 표현하려는 Token 위치만 1 , 나머지는 모두 0 으로 채운다. import torch import torch.nn.functional as F token_id = torch.tensor(2) vocab_size = 5 one_hot = F.one_hot(token_id, num_classes=vocab_size) print(one_hot) print(one_hot.shape) 실행 결과: tensor([0, 0, 1, 0, 0]) torch.Size([5]) Vocabulary 크기가 10,000이면 Token 하나를 나타내기 위해 길이가 10,000인 벡터가 필요하다. 그중 대부분이 0 이므로 이러한 벡터를 희소 벡터(Sparse Vector) 라고 한다. Bag of Words Bag of Words(BOW)는 문서에 각 단어가 몇 번 등장했는지를 벡터로 표현한다. Vocabulary: ["I", "love", "Python"] 문장: "I love Python Python" BOW: [1, 1, 2] 단어 빈도를 간단하게 표현할 수 있지만 다음 두 문장을 구별하기 어렵다. 나는 너를 좋아한다 너는 나를 좋아한다 두 문장에 등장하는 단어가 같다면 BOW 벡터도 비슷해진다. 즉, 단어의 순서와 문맥을 반영하지 못한다. TF-IDF TF-IDF(Term Frequency-Inverse Document Frequency)는 단어가 한 문서에서 자주 나타날수록 가중치를 높이고, 전체 문서에서 흔하게 나타날수록 가중치를 낮춘다. TF-IDF = TF × IDF TF: 특정 문서에서 해당 단어가 등장한 빈도다. IDF: 전체 문서에서 드물게 등장할수록 커지는 값이다. 모든 문서에 자주 나타나는 단어보다 특정 문서의 내용을 잘 나타내는 단어에 높은 가중치를 줄 수 있다. 다만 BOW와 마찬가지로 단어 순서와 문맥을 직접 학습하지는 않는다. 2. Word Embedding Word Embedding은 Token을 저차원의 밀집 벡터(Dense Vector) 로 표현하는 방법이다. Token ID ↓ Embedding Layer ↓ [0.21, -0.72, 0.13, 0.84, ...] 원-핫 벡터는 각 단어가 서로 독립적이지만, Embedding은 학습을 통해 비슷한 문맥에서 사용되는 단어가 가까운 벡터를 갖도록 만들 수 있다. nn.Embedding PyTorch의 nn.Embedding 은 Token ID에 대응하는 벡터를 Embedding Table에서 가져오는 층이다. import torch.nn as nn embedding = nn.Embedding( num_embeddings=10000, embedding_dim=128, padding_idx=0, ) print(embedding) print(embedding.weight.shape) 실행 결과: Embedding(10000, 128, padding_idx=0) torch.Size([10000, 128]) 인자 의미 num_embeddings Vocabulary에 포함된 Token 개수다. embedding_dim Token 하나를 나타낼 벡터 차원이다. padding_idx Padding Token의 ID다. 해당 행은 학습에서 갱신하지 않는다. 입력과 출력 Shape는 다음처럼 변한다. input_ids = torch.tensor([ [1, 2, 3], [4, 2, 5], ], dtype=torch.long) embedding = nn.Embedding( num_embeddings=10, embedding_dim=4, padding_idx=0, ) output = embedding(input_ids) print("입력 Shape:", input_ids.shape) print("출력 Shape:", output.shape) print("학습 여부:", embedding.weight.requires_grad) print("Parameter 수:", embedding.weight.numel()) 실행 결과: 입력 Shape: torch.Size([2, 3]) 출력 Shape: torch.Size([2, 3, 4]) 학습 여부: True Parameter 수: 40 입력 Shape [batch_size, sequence_length] 의 뒤에 embedding_dim 이 추가된다. [2, 3] → [2, 3, 4] Embedding 벡터의 의미를 사람이 직접 지정하는 것은 아니다. 처음에는 초기값으로 시작하고, 모델의 다른 가중치와 함께 학습된다. 3. Word2Vec Word2Vec은 주변 문맥을 이용해 단어의 벡터를 학습한다. 비슷한 문맥에서 자주 사용되는 단어가 벡터 공간에서도 가까워지도록 만든다. CBOW와 Skip-gram 방식 입력 예측 대상 특징 CBOW 주변 단어 중심 단어 비교적 빠르고 자주 등장하는 단어 학습에 유리하다. Skip-gram 중심 단어 주변 단어 계산량은 많지만 드문 단어 학습에 유리하다. 문장이 다음과 같다고 가정한다. 나는 오늘 공원에서 귀여운 강아지를 산책시키며 행복한 시간을 보냈다. 중심 단어가 강아지를 이고 Window 크기가 2라면 주변 문맥은 다음과 같다. ["공원에서", "귀여운", "산책시키며", "행복한"] CBOW는 주변 네 단어로 강아지를 을 예측한다. Skip-gram은 강아지를 로 주변 네 단어를 예측한다. SGNS SGNS(Skip-Gram with Negative Sampling)는 Skip-gram의 계산량을 줄이는 방법이다. 실제 중심 단어와 주변 단어의 조합은 관련이 있는 Positive Sample로 학습한다. 무작위로 선택한 단어 조합은 관련이 없는 Negative Sample로 학습한다. Vocabulary 전체에 대한 확률을 매번 계산하지 않고 일부 단어만 비교한다. NSMC 데이터 전처리 한국어 Word2Vec 학습을 위해 네이버 영화 리뷰 데이터(NSMC)를 사용한다. 노트북의 macOS 환경에서는 MeCab을 다음과 같이 연결했다. import os from konlpy.tag import Mecab os.environ["MECABRC"] = "/opt/homebrew/etc/mecabrc" mecab = Mecab( dicpath="/opt/homebrew/lib/mecab/dic/mecab-ko-dic", ) MECABRC 와 사전 경로는 설치 방법과 운영체제에 따라 달라진다. import urllib.request import pandas as pd urllib.request.urlretrieve( "https://raw.githubusercontent.com/e9t/nsmc/master/ratings.txt", filename="ratings.txt", ) train_data = pd.read_table("ratings.txt") train_data = train_data[:20000] 한글과 공백을 제외한 문자를 제거한다. train_data["document"] = train_data["document"].str.replace( "[^ᄀ-하-ᅵ가-힣 ]", "", regex=True, ) MeCab으로 형태소를 분리하고 불용어를 제거한다. stopwords = [ "도", "는", "다", "의", "가", "이", "은", "한", "에", "하", "고", "을", "를", "인", "듯", "과", "와", "네", "들", "지", "임", "게", ] tokenized_data = [] for sentence in train_data["document"]: tokens = mecab.morphs(sentence) tokens = [word for word in tokens if word not in stopwords] tokenized_data.append(tokens) MeCab은 붙어 있는 한국어 문장을 형태소 단위로 나눈다. mecab.morphs("아버지가방에들어가신다") ['아버지', '가', '방', '에', '들어가', '신다'] Gensim으로 Word2Vec 학습하기 from gensim.models import Word2Vec model = Word2Vec( sentences=tokenized_data, vector_size=100, window=5, min_count=5, workers=2, sg=1, negative=5, ) 인자 의미 vector_size 단어 벡터 차원이다. window 중심 단어 앞뒤로 확인할 단어 범위다. min_count 이 횟수 이상 등장한 단어만 학습한다. workers 학습에 사용할 CPU Worker 수다. sg 0 은 CBOW, 1 은 Skip-gram이다. negative Negative Sampling에 사용할 단어 수다. 학습 결과는 3,990개 단어를 각각 100차원 벡터로 표현했다. print(model.wv.vectors.shape) (3990, 100) most_similar() 는 코사인 유사도를 기준으로 가까운 단어를 찾는다. model.wv.most_similar("블록버스터", topn=5) 데이터를 20,000개 리뷰로 제한했기 때문에 결과 품질은 전체 데이터와 학습 설정에 따라 달라질 수 있다. 4. FastText와 한글 자모 분해 Word2Vec은 Vocabulary에 없는 단어의 벡터를 바로 만들기 어렵다. FastText는 단어를 문자 n-gram 단위로 나누어 학습하므로, 처음 보는 단어나 오타도 이미 학습한 부분 문자열을 조합해 표현할 수 있다. playing → <pl, pla, lay, ayi, yin, ing, ng> 한국어는 조사와 어미가 붙어 단어 형태가 다양하다. 자모 단위로 분해하면 비슷한 철자를 가진 단어와 오타가 더 많은 부분을 공유한다. hgtk 로 한글 분해와 조합하기 import hgtk print(hgtk.letter.decompose("김")) print(hgtk.letter.compose("ᄀ", "ᅵ", "ᄆ")) 실행 결과: ('ᄀ', 'ᅵ', 'ᄆ') 김 받침이 없는 글자는 빈 문자열 대신 - 를 넣어 항상 세 글자 단위로 표현한다. def word_to_jamo(token): def to_special_token(jamo): return jamo if jamo else "-" decomposed_token = "" for char in token: try: cho, jung, jong = hgtk.letter.decompose(char) decomposed_token += ( to_special_token(cho) + to_special_token(jung) + to_special_token(jong) ) except Exception as error: if type(error).__name__ == "NotHangulException": decomposed_token += char return decomposed_token print(word_to_jamo("남동생")) print(word_to_jamo("여동생")) 나ᄆ도ᄋ새ᄋ 여-도ᄋ새ᄋ MeCab 형태소 분석과 자모 분해를 연결한다. def tokenize_by_jamo(sentence): return [word_to_jamo(token) for token in mecab.morphs(sentence)] FastText 실습에는 네이버 쇼핑 리뷰 20만 건을 사용했다. urllib.request.urlretrieve( "https://raw.githubusercontent.com/bab2min/corpus/master/sentiment/naver_shopping.txt", filename="ratings_total.txt", ) total_data = pd.read_table( "ratings_total.txt", names=["ratings", "reviews"], ) 각 리뷰를 형태소로 나눈 뒤 자모 단위로 변환한다. from tqdm import tqdm tokenized_data = [] for review in tqdm(total_data["reviews"].to_list()): tokenized_sample = tokenize_by_jamo(review) tokenized_data.append(tokenized_sample) 분해한 Token을 텍스트 파일로 저장하고 FastText를 학습한다. import fasttext with open("tokenized_data.txt", "w", encoding="utf8") as output_file: for line in tokenized_data: output_file.write(" ".join(line) + "\n") model = fasttext.train_unsupervised( "tokenized_data.txt", model="cbow", ) model.save_model("fasttext.bin") 노트북에서는 남동쉥 , 남동셍ᄏ , 제품^^ 처럼 변형되거나 오타가 있는 입력에서도 원래 단어와 관련된 결과가 나타나는지 확인했다. 남동쉥 → 남동생, 남친, 남매, 남짓, 남녀 ... 제품^^ → 제품, 제풍, 반제품, 완제품, 상품 ... 이는 FastText가 단어 전체만 외우지 않고 문자 조각을 함께 학습하기 때문이다. 5. 시퀀스 데이터 시퀀스 데이터(Sequence Data)는 값의 순서와 앞뒤 관계가 중요한 데이터 다. 데이터 순서가 중요한 이유 자연어 문장 단어 순서가 달라지면 문장의 의미가 달라진다. 주가·날씨 이전 시점의 변화가 다음 시점과 연결된다. 음성 시간에 따른 파형의 흐름이 의미를 만든다. 영상 연속된 Frame의 변화가 움직임을 나타낸다. 일반적인 MLP는 각 입력을 독립적으로 처리한다. 반면 RNN은 이전 시점의 정보를 Hidden State에 담아 다음 시점 계산에 전달한다. 6. RNN의 동작 원리 RNN(Recurrent Neural Network)은 현재 입력과 이전 Hidden State를 함께 사용해 새로운 Hidden State를 만든다. h_t = tanh(W_xh x_t + W_hh h_(t-1) + b) 기호 의미 x_t 현재 시점의 입력이다. h_(t-1) 이전 시점까지의 정보를 담은 Hidden State다. h_t 현재 입력까지 반영한 새로운 Hidden State다. W_xh , W_hh 학습되는 가중치다. tanh 값을 -1~1 범위로 변환하는 활성화 함수다. Hidden State는 가중치 자체가 아니라, 현재 시점까지 본 정보 중 다음 계산에 필요한 내용을 압축한 작업 메모리 다. RNN은 시점마다 새로운 Cell을 학습하는 것이 아니다. 같은 RNN Cell과 같은 가중치를 전체 시간축에서 반복해서 사용한다. 입력 Shape batch_first=True 인 RNN의 입력 Shape는 다음과 같다. [batch_size, sequence_length, input_size] 예를 들어 [32, 20, 10] 은 다음을 의미한다. 한 번에 32개 시퀀스를 처리한다. 시퀀스 하나는 연속된 20개 시점으로 구성된다. 각 시점은 10개 Feature를 가진다. 장기 의존성 문제 기본 RNN은 가까운 과거의 패턴을 처리하는 데 유용하지만, 오래전 정보를 마지막까지 유지하기 어렵다. 같은 가중치를 시간축에서 반복 적용하며 역전파하기 때문에 긴 시퀀스에서는 기울기 소실 또는 기울기 폭주가 발생하기 쉽다. 이 문제를 완화하기 위해 LSTM과 GRU 같은 구조가 사용된다. 7. RNN으로 다음 KOSPI 종가 예측하기 최근 20거래일의 정보를 입력으로 사용해 다음 거래일의 종가를 예측한다. 1일차 입력 + 초기 Hidden State → h1 2일차 입력 + h1 → h2 ... 20일차 입력 + h19 → h20 h20 → Linear Layer → 21일차 종가 예측 데이터 확인과 정제 원본 데이터는 4,513행이며 다음 열을 가진다. 열 의미 Open 시가 High 고가 Low 저가 Close 종가 Adj Close 배당과 주식 분할 등을 반영한 수정 종가 Volume 거래량 날짜를 datetime 으로 바꾸고 필수 열의 결측치를 제거한 뒤 날짜순으로 정렬한다. df["Date"] = pd.to_datetime( df["Date"], dayfirst=True, errors="coerce", ) numeric_cols = ["Open", "High", "Low", "Close", "Adj Close", "Volume"] for column in numeric_cols: df[column] = pd.to_numeric(df[column], errors="coerce") required_cols = ["Date", "Open", "High", "Low", "Close"] df_copy = df.dropna(subset=required_cols).copy() df_copy = df_copy.sort_values("Date").reset_index(drop=True) 정제 후 2000년 1월 4일부터 2018년 1월 22일까지 총 4,452행이 남았다. Feature Engineering 당일 가격과 최근 가격 흐름을 함께 표현하도록 파생변수를 만든다. data = df_copy.copy() data["Return_1d"] = data["Close"].pct_change() data["Range_pct"] = (data["High"] - data["Low"]) / data["Close"] data["OC_pct"] = (data["Close"] - data["Open"]) / data["Open"] data["MA5"] = data["Close"].rolling(5).mean() data["MA20"] = data["Close"].rolling(20).mean() data["MA5_ratio"] = data["Close"] / data["MA5"] data["MA20_ratio"] = data["Close"] / data["MA20"] data["Volatility_10"] = data["Return_1d"].rolling(10).std() data["Target_Close"] = data["Close"].shift(-1) data["Target_Date"] = data["Date"].shift(-1) 사용하는 Feature는 총 10개다. FEATURES = [ "Open", "High", "Low", "Close", "Return_1d", "Range_pct", "OC_pct", "MA5_ratio", "MA20_ratio", "Volatility_10", ] shift(-1) 을 사용하면 현재 행의 Target에 다음 거래일 종가가 들어간다. 이동평균과 다음 날 Target을 만드는 과정에서 생긴 결측치를 제거한 결과 학습에 사용할 데이터는 4,432행이다. 8. 시계열 데이터 분리와 Scaling 시계열 데이터를 무작위로 나누면 미래 데이터가 Train에 포함되고 과거 데이터가 Test에 포함될 수 있다. 따라서 날짜순으로 다음과 같이 분리한다. n = len(data) train_end = int(n * 0.70) val_end = int(n * 0.85) 분할 행 개수 Target 기간 Train 3,102 2000-02-01 ~ 2012-08-20 Validation 665 2012-08-21 ~ 2015-05-04 Test 665 2015-05-06 ~ 2018-01-22
TS. Lê Thị Ánh là chuyên gia trong lĩnh vực kế toán – thuế và quản trị tài chính doanh nghiệp, hiện là Cố vấn chuyên môn tại VFDI. Với nền tảng chuyên môn về kế toán, thuế và tài chính doanh nghiệp, TS. Lê Thị Ánh tham gia tư vấn, xây dựng và định hướng các giải pháp liên quan đến hệ thống kế toán – thuế, quản trị tài chính, kiểm soát số liệu và phòng ngừa rủi ro cho doanh nghiệp. Tại VFDI, TS. Lê Thị Ánh tham gia định hướng chuyên môn cho các nội dung và giải pháp về kế toán – thuế, xây dựng hệ thống kế toán doanh nghiệp, hoàn thiện sổ sách và quyết toán thuế, tư vấn đồng hành kế toán doanh nghiệp, kế toán hộ kinh doanh và kế toán sàn thương mại điện tử. Định hướng chuyên môn tập trung vào sự kết nối giữa Kế toán – Thuế – Tài chính, nhằm nâng cao tính minh bạch, khả năng kiểm soát và giá trị của dữ liệu kế toán đối với hoạt động quản trị doanh nghiệp. VFDI hiện cũng xác định “xây dựng hệ thống thay vì chỉ cung cấp dịch vụ”, kết nối kế toán – thuế – tài chính và chuyên sâu kế toán sàn thương mại điện tử là những điểm khác biệt của thương hiệu. LIÊN HỆ: Địa chỉ: Chung cư Goldsilk Tổ dân phố 7, Phường Hà Đông, Hà Nội Số điện thoại: 0904848855 Website: https://vfdi.vn/nhan-su/ts-le-thi-anh Facebook: https://www.facebook.com/leanheducation Twitter: https://twitter.com/lethianhvfdi Pinterest: https://www.pinterest.com/lethianhvfdi/ Tumblr: https://lethianhvfdi.tumblr.com/ Youtube: https://www.youtube.com/channel/UCzTXsk0YyZvE1Eaaj5ecGwg