Brazil’s Digital Banking Scale Is No Longer the Finish Line: Ken Research Maps the Shift Toward Monetizing Payments, Data and Financial Relationships Brazil’s digital-finance story is entering a more demanding phase. The latest Ken Research framework values the Brazil Digital Banking and Open Finance Market at USD 15,000 million in 2025 and projects it to reach USD 33,978 million by 2031 , representing a forecast CAGR of 14.60% . The commercial question is no longer whether customers will bank digitally; it is how institutions convert digital activity into deeper, risk-adjusted revenue. That distinction matters because Brazil already has the transaction infrastructure and behavioral scale that many emerging digital-finance markets are still trying to build. FEBRABAN reported 240.8 billion banking transactions in 2025 , with 83% occurring through digital channels and 78% through mobile banking. FEBRABAN's Banking Technology Survey therefore points to a market where the smartphone has become the primary banking interface rather than an alternative channel. The counter-thesis is that digital ubiquity does not automatically create attractive economics. Payments can commoditize, fraud and cybersecurity costs can rise with transaction frequency, and aggressive credit expansion can destroy value when funding and loss economics are weak. This tension is visible in the Brazil FinTech Online Lending and Credit Platforms Market , where competition is increasingly shifting toward secured products, underwriting quality, repeat borrowing and risk-adjusted monetization rather than customer acquisition alone. The Market Has Moved Beyond Digital Access The historical phase was about moving financial activity from branches and conventional channels into apps, digital accounts and instant payments. The next phase is structurally different. Ken Research estimates that the market expanded at a historical CAGR of approximately 19.33% during 2020–2025 , but forecasts a more normalized 14.60% CAGR during 2026–2031 . Slower percentage growth does not imply a weaker opportunity; it signals that the economic engine is changing. The report models digital banking transaction volume at approximately 199.9 billion transactions in 2025 , rising to around 445.0 billion by 2031 . Yet transaction count alone will not determine who captures the incremental value. More important will be revenue per active relationship, deposit depth, credit quality, merchant monetization, investment distribution and the ability to turn permissioned financial data into relevant offers. Brazil’s next digital-banking advantage will come less from putting another account on a phone and more from becoming the institution that customers repeatedly use to move, store, borrow, invest and manage money. Pix Has Made Frequency Abundant—Now Banks Must Monetize It Pix has effectively transformed payment frequency into shared infrastructure. The Banco Central do Brasil reports more than 170 million individual Pix users, equivalent to about 80% of the population, and more than 7 billion transactions in May 2026 . When a payment rail reaches that degree of penetration, basic money movement becomes increasingly difficult to defend as a standalone differentiator. The strategic implication is visible across the broader Brazil Retail Banking Market : frequent digital interactions can become an acquisition and retention engine for deposits, cards, lending, insurance and investments, but only when institutions have the product breadth and analytics to act on those interactions. Payment engagement is therefore the top of a commercial funnel, not the final product. Where Payment Frequency Can Create Higher-Value Economics Deposits: frequent account activity strengthens the opportunity to become the customer's primary liquidity account. Credit: transaction histories can improve affordability assessment, offer timing and behavioral risk segmentation. Merchant services: payment flows can support acquiring, working-capital products and merchant cash-flow analytics. Investments: surplus-balance visibility creates opportunities for automated savings and investment distribution. Retention: recurrent financial routines raise switching friction even when basic payment rails remain interoperable. The economic challenge is that high payment frequency also increases the operational cost of poor fraud controls, weak authentication or unreliable infrastructure. Scale therefore rewards institutions that can process enormous volumes at low unit cost while keeping customer friction, losses and downtime under control. Consent Is Becoming a Commercial Asset, Not Just a Compliance Workflow Open Finance adds a second strategic layer to Brazil’s digital infrastructure: permissioned portability of financial information and services. The Banco Central do Brasil's Open Finance guidance explains that customers can authorize sharing of account, card, credit, investment and other financial information between participating institutions, while payment services can also be initiated through selected third-party journeys. Ken Research records approximately 103 million active Open Finance consents in September 2025 and approximately 68 million connected accounts. That scale changes the competitive meaning of customer data. Incumbents can no longer assume that historic account ownership gives them an exclusive informational advantage, while challengers can increasingly use consented histories to improve aggregation, underwriting and personalized product distribution. The Value Is in Conversion, Not Consent Count A consent has little commercial value if the receiving institution cannot transform it into a better customer outcome. The more important operating questions are whether connected data improves credit approval quality, raises product conversion, reduces manual onboarding, supports smarter pricing or helps users consolidate financial activity inside one preferred interface. Personalized underwriting: richer cash-flow histories can improve credit selection beyond static bureau information. Account aggregation: consolidated financial views can increase app utility and engagement. Product comparison: portability can make rates, fees and service quality more transparent. Payment initiation: Open Finance can move beyond data sharing into transaction origination. Customer recovery: institutions can target refinancing or relationship-deepening opportunities using permissioned data. Distribution Is Escaping the Bank-Owned Interface Mobile banking apps remain the principal customer interface, but the distribution architecture is widening. The primary report expects Open Finance API channels and embedded-finance channels to expand faster than conventional web banking as financial services appear inside merchant software, marketplaces and third-party digital experiences. That moves competition from “who owns the banking app?” toward “who can place the right financial product inside the right customer journey?” This same platform logic is evident in the Brazil Digital Wallet & Superapps Ecosystem Market , where interoperable payments are increasingly layered with merchant services, credit, investments and wider financial functionality. For banks and fintechs, API distribution can lower dependence on proprietary traffic, but it also exposes products to more direct comparison and makes integration quality a commercial capability. APIs Change Customer-Acquisition Economics An embedded credit offer inside merchant software, a payment initiated from another institution’s interface or an investment recommendation triggered by aggregated balances can reach customers without requiring them to begin their journey inside the provider's own application. That can reduce acquisition friction, but it also shifts value toward institutions that combine reliable APIs, fast decisioning, strong partner management and disciplined economics. Competition Is Shifting From Account Scale to Relationship Quality The primary report identifies Nubank, Itaú Unibanco, Banco do Brasil, Banco Inter and Mercado Pago among the major participants, within a much broader ecosystem of universal banks, digital-only banks, payment institutions and fintech/API providers. The market should not be read as a simple incumbent-versus-challenger contest. Each institution enters the next phase with different strengths in funding, customer acquisition, deposits, merchant reach, credit, data, technology and distribution. The competitive variables that matter most are therefore becoming more operational. Low-cost customer acquisition is valuable only if users become active; high transaction activity is valuable only if it supports profitable relationships; rich data is valuable only if models can convert it into better decisions; and broad product catalogues matter only if cross-selling improves lifetime economics. Engagement quality: frequency and depth of meaningful financial activity. Funding economics: ability to support credit without excessive balance-sheet cost. Risk-adjusted credit: growth that survives delinquency and collection costs. API reliability: dependable integration with Open Finance and embedded partners. Fraud control: prevention and recovery without creating excessive customer friction. Product depth: ability to monetize payments through deposits, credit, investment and merchant relationships. The Biggest Constraints Are Now Inside the Operating Model The same infrastructure that creates opportunity also raises execution standards. Open Finance requires consent management, secure data processing and reliable APIs across a large network of regulated institutions. Pix creates extreme transaction frequency and correspondingly high expectations for uptime, fraud monitoring and rapid incident response. Digital credit adds another layer of exposure through funding costs, affordability and portfolio quality. Regulation is also evolving from framework creation
Brazil’s Farms Are Moving From Connected Tools to Digital Operating Systems: Ken Research Maps the Next Agritech Value Pool Brazil’s agritech platforms and smart farming market is entering a different phase of digitalization. Proprietary estimates from Ken Research value the market at USD 594 million in 2025 and project it to reach USD 1,221 million by 2031 , implying a 12.76% CAGR. The commercial story is therefore becoming less about whether farms use digital tools and more about how deeply software, data and automation become embedded in everyday farm decisions. The mechanism behind that growth matters. Remote sensing, connected machinery, farm-management software, artificial intelligence and precision agronomy are increasingly being combined rather than purchased as isolated technologies. Ken Research estimates that digitally managed agricultural area could expand from 23.7 million hectares in 2025 to 45.8 million hectares in 2031 , creating a larger base over which providers can monetize recurring analytics, agronomic intelligence and operational services. The counter-thesis is equally important: Brazil’s enormous agricultural scale does not guarantee uniform technology adoption. Connectivity, financing, interoperability and affordability still constrain deployment, particularly outside large commercial farms. That distinction becomes clearer when digital agriculture is viewed beside the Brazil Agriculture Market , where overall farm economics are driven by very different producer sizes, crop systems, regions and capital structures. The market scope covers provider revenue from farm-management platforms, digital agronomy, crop intelligence, remote sensing, connected field systems and related smart-farming services. Conventional agricultural machinery without a meaningful digital solution component sits outside the core revenue lens. Brazil’s Agricultural Scale Makes Small Efficiency Gains Valuable Digital agriculture becomes economically compelling when a relatively small operational improvement can be multiplied across a very large production footprint. Brazil provides precisely that environment. According to Companhia Nacional de Abastecimento , the 2024/25 grain crop covered 81.7 million hectares and produced an estimated 350.2 million tonnes . Conab also reported a 13.7% increase in average crop productivity for that season. Smart-farming providers do not create that national productivity result themselves, but the production scale shows why technologies that improve input timing, planting decisions, machinery utilization or loss prevention can carry meaningful economic value. A small per-hectare improvement can become material when applied across large soybean, corn, cotton and other commercial crop operations. The Revenue Model Can Deepen Without Equivalent Land Expansion The more important strategic shift is therefore from selling a first digital tool to monetizing more decisions on the same farm. A platform that begins with field records or satellite imagery can potentially extend into prescriptions, machinery data, weather intelligence, crop monitoring, financial planning and traceability. That expands revenue per connected hectare without requiring agricultural land to grow at the same rate. Farm-management software: consolidates agronomic, operating and financial workflows. Remote sensing: expands field visibility without proportional physical inspection costs. AI-supported agronomy: converts multiple datasets into recommendations and alerts. Connected machinery: links equipment performance to planting, spraying and harvesting decisions. Compliance data: turns field records into auditable information for buyers, lenders and other stakeholders. From Observation to Intervention: Platforms Are Moving Deeper Into Farm Operations Early agricultural digitization often improved visibility: where a machine was, what weather was approaching or how a field looked from above. The higher-value opportunity is intervention. Platforms become commercially harder to replace when they influence what happens next—when to irrigate, which field requires attention, how inputs should be allocated or whether machinery is operating efficiently. This shift overlaps with the Brazil Agri Drones and Precision Agriculture Market , where aerial data collection and precision-farming technologies expand the information available for field-level decisions. The strategic value lies less in the sensor or aircraft alone than in converting collected data into an action that improves farm economics. Ken Research identifies AI-enabled farm-management and decision platforms as a central solution layer in the current market. Commercial-farm platform penetration is estimated at 34% in 2025 and modeled to reach 58% by 2031 . That still leaves adoption headroom, but future conversion will depend increasingly on demonstrable return rather than technology novelty. Supply Is Deep, but Integration Is Becoming the Real Differentiator Brazil does not lack agritech suppliers. The Radar Agtech Brasil 2024 research associated with Embrapa mapped roughly 1,972 agtechs across the agricultural value chain. Of these, 818 operated in categories inside the farm, illustrating the density of technology supply around management, automation, monitoring and related production activities. That abundance creates choice, but it also creates integration friction. Farms can accumulate separate applications for weather, agronomy, machinery, financial management, imagery and compliance. The resulting problem is not simply “too many apps”; it is fragmented operational data, duplicate workflows and difficulty establishing one reliable decision layer across the enterprise. Interoperability Changes Retention Economics The next competitive battleground is likely to involve how easily platforms connect with equipment, external data sources and other applications. The primary market assessment identifies 157 property-management startups and 144 data-integration platforms within the mapped 2024 ecosystem. That density creates pressure for open interfaces and useful integrations rather than closed data silos. Data coverage: how much of the farm workflow can the platform see? Interoperability: can it ingest information from multiple machines and systems? Agronomic intelligence: does data translate into actionable recommendations? Workflow depth: is the software used for occasional monitoring or everyday decisions? Recurring value: does the customer gain measurable benefit every season? Providers including Solinftec, Agrosmart, Cropwise, Climate FieldView and Farmbox illustrate the varied competitive approaches present in the market. The primary report does not publish verified individual market-share percentages for these companies, so they should be viewed as an unranked participant set rather than a league table. The Platform Opportunity Extends Beyond Agronomy Once reliable farm data exists, its usefulness extends outside field operations. Procurement, lending, insurance, trading and supply-chain relationships can all benefit from better digital records. This creates a broader platform opportunity in which agronomic information can become part of the commercial infrastructure surrounding the producer. That connection is visible in the Brazil Digital Agriculture Marketplaces Market . These marketplaces connect producers with input suppliers, equipment providers, commodity buyers, financial institutions and service providers, demonstrating how digital workflows increasingly bridge production decisions and transactions rather than treating them as separate ecosystems. The implication for smart-farming vendors is significant. A provider that controls a trusted decision layer may have opportunities to integrate services adjacent to production, but expansion also increases complexity. Handling agronomic records is different from supporting financial decisions, transactions or supply-chain verification, and providers must avoid adding modules that weaken the usability of their core product. Traceability Turns Farm Data Into Commercial Infrastructure Brazil’s export position adds another source of technology demand. The Ministério da Agricultura e Pecuária reported agribusiness exports of USD 169.2 billion in 2025 , equivalent to 48.5% of Brazil’s total export value. An export-oriented production system increases the commercial importance of knowing where products came from, how fields were managed and whether required information can be demonstrated to counterparties. This does not mean every farm needs the same compliance technology, but it expands the potential buyer set for geospatial monitoring, traceability and auditable farm records beyond growers themselves. Data Quality Becomes a Product Attribute As farm information supports decisions by processors, traders, financial institutions or other third parties, accuracy and continuity become economically relevant. A platform that captures incomplete or incompatible records may be adequate for basic monitoring but less useful when information must support formal commercial workflows. This changes product development priorities. Data governance, API reliability, historical continuity and verification mechanisms become part of the value proposition alongside dashboards and analytics. In effect, the market begins to reward platforms not only for generating insights but also for maintaining trustworthy digital records. The Biggest Adoption Risk Is Uneven Farm Economics Brazil’s agricultural sector contains both extremely large, technology-intensive commercial farms and smaller operations with very different capital and service needs. A solution that produces clear payback across thousands of hectares may be difficult to justify on a smaller farm when subscription fees, hardware requirements, connectivity or implementation effort are spread over fewer hectares. This makes total cost of adoption more important than headline software
오늘은 오랜만에 VM을 설치하는 일을 했습니다 예전 가상머신 설치 때도 디스크 파일이 깨져 애 먹었던 기억이 있는데 역시나.. 설치하라는대로 했지만 또 디스크 파일이 깨졌다네요 오늘은 이 사례를 트러블슈팅 남겨보겠습니다 체크 상황1. 디스크 용량이 큰가? 디스크 용량은 14기가 정도 되었습니다 크기가 너무 커서 문제인가라고 생각하기에는 용량이 범위를 크게 벗어나진 않았습니다 체크 상황2. 체크섬 비교 1. SHA-256으로 검증하기 파워쉘에 Get-FileHash "C:\Users\admin\Downloads\파일이름.iso" -Algorithm SHA256 명령어를 통해 해시 값을 얻습니다 이때 14GB의 용량이기에 좀 오래걸려서 결과값이 바로 나오지 않는데 1-3분 기다리면 나옵니다 나온 해시값을 비교해봤는데 문제가 없었습니다 그래서 그냥 제거하고 다시 만들었어요^^ 이전과의 차별점 가상머신을 만들때 ** proceed with Unattended Installation** 이걸 체크해서 설정을 하나하나 다 했었는데 이번엔 체크를 해제했다. 그랬더니 별 설정을 건들지 않고 바로 설치가 됐다 이제 실행해보자~ 근데.. 이번엔 디스크 파일이 깨진건 아닌데 시작을 하면 중단이 되는 상황이;;; 발생했다 문제 상황 1. 전원 켜는 중 E_FAIL (0x80004005) 전원 켜는 중 이라는 메시지가 떴고 100%도 됐는데 중단이라고 떠있고 로그도 별다른게 없었다. 여기서 시간을 진짜 그냥 영원히 썼다 증상 : VM 시작 시 "전원 켜는 중 100%"까지 갔다가 중단되고, 결과 코드에 E_FAIL (0x80004005) 가 떴다. 오류 창에 자세히(Details) 버튼도 없어서 뭐가 문제인건지 원; 용어 : E_FAIL 은 "실패했다"는 뜻의 포괄적인 코드라서, 이것만으로는 원인을 알 수 없다. 원인은 로그에서 찾아야 한다 시도한 것 (Hyper-V 충돌 가설) optionalfeatures 에서 Hyper-V 등 기능 확인 ** bcdedit /set hypervisorlaunchtype off** 후 재부팅 → 증상 동일 되돌릴 때는 bcdedit /set hypervisorlaunchtype auto 로그 확인 VM 폴더의 Logs\VBox.log를 확인했다 Get-Content "$env:USERPROFILE\VirtualBox VMs\[가상머신]\Logs\VBox.log" 로그가 1.4KB로 매우 짧았다. 정상 실행이면 수백 KB가 되므로, 부팅 전에 실패했다는 뜻이다 Machine state changed to 'Starting' 직후 Power up failed (... E_FAIL)이 찍혔다 VERR_NEM, VT-x 같은 가상화 관련 문구는 없었다 로그에 Core Isolation (Memory Integrity): ENABLED가 있었지만 이것이 원인이라는 증거는 없었다 원인의 주범 발견 저장소 확인 저장소를 확인해봤다 가상머신을 선택한 뒤 설정->저장소에서 가상머신.vdi 에 경고 표시가 있었다. 이전에 지운 VM과 등록 정보가 꼬였을 가능성이 높았다 해결 기존의 가상머신을 삭제하고, 이름을 바꿔서 새로 만들었더니 정상적으로 시작되었다 그리고 .vdi 파일도 제거했다가 다시 설정해줬다 문제를 해결하며.. 사실 정확한 원인은 확정하지 못했다. VM 삭제와 재생성을 하며 남은 디스크/등록 정보의 충돌로 추정한다. 이 문제에서는 Hyper-V 설정이나 메모리 무결성이 원인은 아니었다. 예전에도 비슷한 일을 겪었었는데 보통 이런 문제였던 것 같다. 뭔가 쎄하면 그냥 클린하게 제거 후 재생성해도 좋을 것 같다 이번에 배운 점 오류는 로그 파일의 맨 아래에서 위로 읽는다 E_FAIL 같은 포괄적 코드는 원인이 아니므로 로그의 VERR_ 문구를 찾는다 VBox.log가 아주 작으면 부팅 전 실패다 VM을 삭제했다 다시 만들 때는 "모든 파일 삭제" 를 선택하고, 막히면 이름을 바꿔서 만든다 오늘도 끝!!!!
4.2 LSTM과 GRU 가장 단수한 형태의 RNN (Vanilla RNN) -> 한계 극복 : LSTM 1. 바닐라 RNN 한계 바닐라 RNN은 가장 기본이 되는 것으로서 출력 결과가 이전 계산 결과에 의존한다. 하지만 비교적 짧은 Sequence에 대해서만 효과를 보인다는 단점으로 시점이 길어질 수록 정보가 앞 뒤로 충분히 전달되지 못하는 현상이 발생한다. -> 만약 다음 내용을 예측학려고 하는데 RNN이 충분한 기억려을 가지지 못한다면 엉뚱한 내요으로 예측하게 된다 : 장기 의존성 문제 2. 바닐라 RNN 내부 2개의 입력(입력값 x_t , 은닉층 값 h_(t-1)) 가중치와 곱해져 메모리 셀의 입력. 하이퍼볼릭탄젠트 함수의 입력으로 사용하고 이 값은 은닉층의 출력인 은닉상태가 된다. 3. LSTM(Long Short-Term Memory) 전통적인 RNN의 단점인 시퀀스가 길어질 수록 정보를 충분히 반영하지 못한다는 단점을 보완한 장단기 메모리이다. 은닉층의 메모리 셀에 입력, 망각, 출력 게이트를 추가하여 불필요한 기억은 지우고, 필요한 기억을 정한다. 셀상태 : C_t (쉽게 말하면 기억 유지 통로) 이 셀 상태는 외쪽에서 오른쪽으로 가는 굵은 선으로 이전의 RNN에서 이전 시점의 은닉상태를 바탕으로 현재 은닉상태를 계산하는 것 처럼 이전 셀의 상태가 다음 셀 상태를 구하기 위한 입력으로 사용 된다. 3개의 게이트 추가 입력 게이트 출력 게이트 삭제 게이트 3-1 입력 게이트 입력게이트 : 현재 정보를 위한 게이트 2가지 값으로 구분됨. 하나는 i_t 시그모이드 함수값(0과 1 사이의 값으로 변환, 활성화 함수), g_t하이퍼볼릭 탄젠트 값(-1과 1사이의 값) 위 2가지 값으로 선택된 기억할 정보의 양을 정하고, 어떻게 결정하는지 활용된다. 3-2 삭제 게이트 삭제 게이트 : 기억을 삭제하기 위한 게이트 여기서는 시그모이드 활성화 함수를 적용한 값만 사용함. 이 갑이 삭제과정을 거친 정보의 양으로 0에 가까울 수록 정보가 많이 삭제 된 것이고, 1에 가까울 수록 정보를 온전히 기억한 것 이 값을 가지고 셀 상태를 구하게 됨. 3-3 셀 상태 여기서 셀 상태는 LSTM의 장기 상태라고 한다. 입력 게이트 구간에서 두가지 값 i_t, g_t 값에 대해서 원소별 곱을 진행. (행렬곱을 하는데 같은 위치에 있는 성분끼리 곱하는 것 -> dot product) = 기억할 값 입력 게이트와 삭제 게이트의 결과값을 더하면 현재 t 시점의 셀 상태라고 하면 t+1 시점의 LSTM 셀로 넘어가짐. 입력 게이트가 0이 되면 삭제 게이트에게만 의존해 C_t 값은 C_t-1 값에만 의존하는 셀 상태값이 됨. 삭제 게이트가 0이 되면 입력 게이트에게만 의존해 C_t 값의 영향력은 0이 되며, 오직 입력 게이트의 값이 현재 셀의 상태값을 결정하게 된다. 삭제 게이트는 이전 시점의 입력이 얼마나 반영할지를 의미하고, 입력게이트느 현재 시점의 입력을 얼마나 반영할지를 결정 3-4 출력 게이트 현재 시점의 시그모이드 함수를 지난 값이고, t t시점의 은닉 상태를 결정한다. -> 은닉 상태를 단기상태라고 하며, 장기상태의 값이 하이퍼볼릭탄젠트 함수 값을 지나 -1 과 1 사이의 값이 된다. 해당 값은 출력 게이트의 값과 연산되며, 값이 걸러지는 효과가 발생함 4. 파이토치의 nn.LSTM() nn.RNN(input_dim, hidden_size, batch_first=True) nn.LSTM(input_dim, hidden_size ,batch_first=True) 기존 RNN 셀 함수를 설정하는 것과 동일한 방식 5. GRU(Gated Recurrent Unit) LSTM의 장기 의존성 문제에 대해 해결책을 유지하면서, 은닉 상태를 업데이트하는 계산은 줄임. -> LSTM을 간단화 시킴 기존 LSTM은 사용하면서 최적의 하이퍼파라미터를 찾아냈다면 GRU 사용할 필요 없음. 둘의 성능 면에서는 차이가 없으나 ᄀ데이터의 양이적을 때에는 GRU가 조금 더 낫고, 데이터의 양이 많다면 LSTM이 더 낫다고 한다. 6.파이토치의 nn.GRU() 코드는 비슷하다. nn.GRU(input_dim, hidden_size, batch_first=True)
1. 컴퓨터 비전 — 분류, 감지, 분할 컴퓨터는 이미지를 어떻게 볼까 우리는 길거리 사진 한 장을 보자마자 사람, 자동차, 버스, 나무를 알아본다. 어디에 있는지도 알고, 도로와 인도의 경계도 구분한다. 그런데 이 사진의 일부를 16×16으로 확대하면 색점 덩어리다. 컴퓨터에게 처음부터 "버스"나 "도로"라는 의미가 주어지는 게 아니다. 컴퓨터가 입력받는 건 각 픽셀의 색을 나타내는 숫자 값 뿐이다. 그래서 컴퓨터 비전 모델은 이 숫자들 사이에서 반복되는 패턴을 학습하고, 그 패턴으로 물체와 장면의 의미를 예측한다. 이렇게 이미지와 영상에서 의미 있는 정보를 이해하고 추론하게 만드는 분야가 컴퓨터 비전 이다. 같은 이미지, 다른 출력 같은 이미지라도 우리가 무엇을 알고 싶은지에 따라 모델이 내야 할 답이 달라진다. 분류(Classification) — 이미지 전체가 무엇인가? → 도심 도로 감지(Detection) — 무엇이 어디에 있는가? → 물체별 박스 + 클래스 분할(Segmentation) — 각 픽셀이 무엇인가? → 픽셀별 클래스 지도 * 분류는 이미지마다, 감지는 물체마다, 분할은 픽셀마다! * CNN과 특징 추출 세 과제의 파이프라인을 나란히 놓으면 재밌는 게 보인다. 분류 : 입력 이미지 → 특징 추출 → 클래스별 점수 감지 : 입력 이미지 → 특징 추출 → 박스·클래스·신뢰도 분할 : 입력 이미지 → 특징 추출 → 픽셀별 클래스 가운데 "특징 추출"이 셋 다 똑같다. 그리고 이 자리에 들어가는 게 2주차에 배운 CNN 이다. 필터를 슬라이딩하면서 얕은 층에서는 선과 모서리를, 깊은 층에서는 눈이나 바퀴 같은 걸 뽑아내는 그 과정 그대로다. 그러니까 세 과제의 차이는 특징을 뽑는 부분이 아니라, 그 뒤에 무엇을 붙이느냐 에 있다. ┌──────────┐ ──→ 분류기 ──→ "고양이" 입력 이미지 ──→ │ CNN │ ──→ 박스 예측기 ──→ 박스 + 82% └──────────┘ ──→ 디코더 ──→ 픽셀 지도 ↑ Backbone 이렇게 공통으로 쓰이는 앞부분을 Backbone(등뼈) 이라고 부른다. "ResNet은 분류 논문이지만 이후 감지·분할 모델의 특징 추출 Backbone으로도 널리 활용된다" 고 한 것이다. 세 모델이 서로 경쟁하는 게 아니라, ResNet이 앞을 맡고 나머지 둘이 뒤를 맡는 쪽에 가까웠다. 이미지 분류 분류는 미리 정해둔 클래스 중 하나를 고르는 과제다. 클래스 종류와 학습 이미지의 정답은 사람이 준비하고, 모델은 클래스를 구분하는 패턴을 학습한다. 골든 리트리버 사진을 넣으면 이렇게 진행된다. CNN이 특징을 추출한다 클래스별 원점수인 로짓(logits) 을 계산한다 로짓은 아직 확률이 아니다. softmax 를 적용하면 확률로 바뀐다 → 86%, 9%, 3%, 2% 가장 높은 걸 최종 결과로 선택한다 로짓을 보면 음수도 있고 다 더해도 1이 아니다. 그래서 확률이라고 부를 수 없다. softmax를 거치면 전부 양수가 되고 합이 100%가 된다. 네 개를 한꺼번에 계산해서 합이 100%가 되게 맞춰주는 장치가 필요한데, 그게 softmax다. 하는 일은 로지스틱 회귀와 똑같고, 선택지 개수만 늘어난 셈이다. ResNet 그럼 성능을 높이려면 CNN을 계속 깊게 만들면 될까? ResNet 논문이 이 질문에서 출발한다. 왼쪽이 훈련 오차, 오른쪽이 테스트 오차다. 20층과 56층을 비교한 건데 둘 다 56층이 더 높다. 테스트 오차만 높았으면 과적합이라고 할 수 있는데, 훈련 데이터에서도 성능이 떨어졌으니 과적합으로는 설명이 안 된다. 논문은 이걸 degradation problem 이라고 불렀다. ResNet은 이 문제를 shortcut connection 으로 해결했다. 기존: 원하는 출력 H(x)를 통째로 학습 ResNet: 입력과 출력의 차이 F(x) = H(x) − x 만 학습 최종 출력: H(x) = F(x) + x 여기서 x는 원본 사진이 아니라 이전 층에서 전달된 특징 맵 이고, F(x)는 필요한 변화량, 즉 Residual 이다. 필요한 변화가 없으면 F(x)를 0에 가깝게 학습해서 입력 x를 그대로 통과시킬 수 있다. 객체 탐지 감지는 무엇이 있는지뿐 아니라 어디에 있는지 도 예측한다. 모델이 특징을 추출한 뒤 각 물체의 위치를 사각형으로 표시하는데, 이걸 Bounding Box 라고 한다. 그리고 검출된 물체마다 세 가지가 함께 나온다. Bounding Box — 물체를 감싸는 사각형의 위치 클래스 — 그게 무엇인지 (dog, bicycle, car) Confidence Score — 그 예측이 얼마나 확실한지 (82%, 91%, 74%) 분류는 이미지 하나에 답이 하나였는데, 감지는 한 이미지에서 여러 물체에 대한 결과 가 나온다. YOLO 대표적인 감지 모델인 YOLO 다. You Only Look Once의 약자다. 이미지 크기 조정 — 네트워크에 맞게 고정 크기로 변환 한 번의 Forward Pass — 하나의 CNN이 전체 이미지를 한 번 처리해서 여러 박스와 클래스를 동시에 예측 NMS — 같은 물체를 가리키는 중복 박스를 정리 YOLO의 핵심은 후보 영역을 따로 만들고 영역별로 분류하는 과정을 나누지 않고, 하나의 네트워크로 직접 예측한다 는 점이고 그래서 빠르다. 2016년에 나온 최초 버전인 YOLOv1은 이미지를 7×7 격자 로 나눈다. 이미지를 여러 칸으로 나눈다 정답 박스의 중심점이 들어 있는 셀이 그 물체를 담당한다 각 셀은 여러 박스 후보와 신뢰도, 클래스 확률을 예측한다 2번에서는 박스 전체가 한 셀 안에 들어갈 필요가 없고, 중심만 있으면 그 셀이 책임진다. --> 여러 셀이 같은 물체를 두고 겹치지 않음. 다만 YOLOv1은 작은 물체나 서로 가까이 있는 물체를 감지하는 데 한계 가 존재함. 이미지 분할 감지는 위치를 사각형으로 표현한다. 그런데 세포처럼 정확한 형태와 경계까지 구분해야 하면 사각형만으로는 부족하다. 이때 쓰는 게 분할 이다. 왼쪽이 현미경 원본이고 오른쪽이 분할 마스크다. 흰색이 세포, 검은색이 배경이다. 감지는 물체를 사각형 으로 찾고, 분할은 물체의 실제 모양 을 픽셀 단위로 구분한다. 그래서 세포의 면적이나 정확한 경계를 분석할 때 유용하다고 한다. U-Net — Encoder–Decoder와 Skip Connection 분할 모델인 U-Net 은 이름 그대로 전체 구조가 알파벳 U자 모양이라 붙은 이름이다. Encoder (왼쪽) — 해상도를 줄이면서 더 넓은 영역의 정보를 모아, 이미지에 무엇이 있는지 판단할 문맥을 학습 Decoder (오른쪽) — 낮아진 해상도를 다시 높여 픽셀별 분할 지도를 만든다 그런데 해상도를 줄이는 과정에서 세밀한 위치와 경계 정보가 약해진다. 이걸 보완하려고 쓰는 게 가로 방향의 Skip Connection 이다. 인코더의 고해상도 특징을 대응하는 디코더 단계로 직접 전달해서 결합한다. Resnet과 U-net의 차이점(우회연결) ResNet은 블록의 입력을 변한 결과에 "더하고", U-Net은 인코더의 특징을 디코더 특징과 "이어 붙인다." 세 모델 정리 모델 ResNet YOLO U-Net 대표 과제 분류 감지 분할 핵심 아이디어 Residual Connection 단일 단계 동시 예측 Encoder–Decoder, Skip Connection 출력 단위 이미지별 클래스 물체별 박스·클래스 픽셀별 클래스 대표 평가 지표 Accuracy, Top-1/Top-5 Error IoU, mAP IoU, Dice 셋 다 CNN 기반이지만 푸는 문제와 출력 형태가 다르다. 그리고 ResNet은 분류 논문이지만 이후 감지·분할 모델의 특징 추출 Backbone으로도 널리 쓰인다 고 한다. 2. NLP 기초 — 임베딩과 RNN 문장을 숫자로 바꾸기 자연어 처리의 첫 단계는 텍스트를 숫자로 바꾸는 것이다. 신경망은 숫자 텐서를 입력으로 처리하기 때문에 텍스트를 수치 표현으로 변환해야 한다. 일반적인 흐름은 세 단계다. 1. 토큰화 — 문장을 처리 단위로 나눔 2. 토큰 ID — 각 토큰에 번호를 부여 3. 벡터 표현 — 숫자 벡터로 변환 원-핫 인코딩과 한계 초기에 많이 쓰던 방식이 원-핫 인코딩 이다. 어휘 집합의 크기를 벡터 차원으로 정하고, 해당 단어의 위치만 1로 표시한다. 어휘가 5개라면 → 고양이 = [0, 1, 0, 0, 0] 단순한데 한계가 세 가지 있다. 차원 증가 — 어휘가 커질수록 벡터도 길어진다 희소성 — 대부분의 값이 0이다 의미 관계 부재 — 비슷한 단어도 서로 다른 값으로만 표현된다 임베딩 이걸 보완하는 게 임베딩 이다. 각 토큰을 고정된 길이의 실수 벡터 로 바꾸는 방법이다. 강아지 [ 0.21, 0.75, −0.18, ...] 고양이 [ 0.24, 0.70, −0.12, ...] 자동차 [−0.61, 0.08, 0.53, ...] 강아지와 고양이는 숫자가 서로 비슷하고, 자동차만 부호까지 다르다. 비슷한 문맥에서 사용한 단어는 벡터 공간에서도 가까워질 수 있다 는 게 이런 뜻이었다. 비교 원-핫 임베딩 벡터 크기 어휘 수만큼 큼 작게 설정 가능 값의 형태 대부분 0 실수로 채워진 밀집 벡터 단어 관계 표현하기 어려움 거리로 비교 가능 한 가지 더 나온 게 정적 임베딩과 문맥적 임베딩 의 차이다. Word2Vec 같은 정적 임베딩은 한 단어에 벡터 하나 만 쓴다. 그래서 여러 의미를 가진 단어를 문맥마다 다르게 표현하기 어렵다. 이후 등장한 문맥적 임베딩은 문장에 따라 표현을 다르게 생성 할 수 있다고 한다. Word2Vec — CBOW와 Skip-gram Word2Vec 은 주변 단어와 중심 단어의 예측 관계를 이용해 단어 임베딩을 학습하는 방법이다. 같은 문장 나는 [영화를] 좋아한다 를 두고 방향이 반대다. CBOW — 주변 단어로 중심 단어를 예측 입력: 나는, 좋아한다 정답: 영화를 Skip-gram — 중심 단어로 주변 단어를 예측 입력: 영화를 정답: 나는, 좋아한다 학습 결과로 비슷한 문맥에서 사용되는 단어는 비슷한 벡터를 갖게 된다. 장점이 인상적이었다. 별도의 라벨링 없이 문장 속 중심 단어와 주변 단어의 관계를 그대로 학습 신호로 쓸 수 있다 는 것이다. 따라서 사람이 정답을 일일이 붙여줄 필요가 없다. 문장에서 순서가 중요한 이유 여기까지는 단어 하나하나를 숫자로 바꾸는 얘기였다. 그런데 문장의 의미를 이해하려면 순서도 중요하다. 내가 빵을 먹었다. 빵이 나를 먹었다. 사용한 단어가 같아도 순서가 바뀌면 의미가 달라진다. 구조 정보 처리 방식 문장 처리에서의 특징 순방향 신경망 입력 → 출력 방향으로 계산 이전 시점의 내부 상태를 다음 시점으로 순환 전달 X RNN 이전 상태를 현재 입력과 함께 사용 앞에서 읽은 정보를 다음 시점으로 전달 RNN과 은닉 상태 RNN 은 현재 입력과 이전 은닉 상태를 결합해 새로운 은닉 상태를 계산한다. 핵심 구조는 세 가지다. 현재 입력 x_t — 현재 시점에 들어오는 단어의 임베딩 이전 상태 h_t−1 — 이전 시점까지 처리한 정보의 요약 새로운 상태 h_t — 현재 단어까지 반영해 다음 시점으로 전달 "이 영화는 정말 재미있다"를 처리한다면, 재미있다 를 읽을 때 그 단어만 보는 게 아니라 앞에서 처리한 이 영화는 정말 의 정보가 은닉 상태를 통해 함께 전달된다. 그리고 각 시점에서 같은 가중치를 공유하되 은닉 상태 값만 달라진다. CNN에서 필터를 공간 방향으로 공유했다면, RNN은 시간 방향으로 공유하는 셈이다. 참고로 은닉 상태가 문장을 글자 그대로 저장하는 건 아니다. 제한된 길이의 벡터 안에 필요한 정보를 압축 해서 담는다. RNN의 입출력 구조 구조 설명 예시 일대일 입력 1개 → 출력 1개 이미지 분류 다대일 입력 시퀀스 → 출력 1개 문장 감성 분석 일대다 입력 1개 → 출력 시퀀스 이미지 캡셔닝 다대다 입력 시퀀스 → 출력 시퀀스 주식 가격 예측 같은 RNN인데 어디서 출력을 뽑느냐만 바꾸면 전혀 다른 과제 가 된다. RNN의 문제점 1. 장기 의존성 학습의 어려움 멀리 떨어진 단어 사이의 관계를 은닉 상태에 오래 유지하기 어렵다. 관련된 학습 문제가 기울기 소실과 폭발 인데, 소실되면 먼 시점의 학습이 어려워지고 폭발하면 학습이 불안정해진다. "나는 아침에 우산을 챙겨 학교에 갔습니다. 오후에 비가 오자 그것 을 펼쳤습니다." 마지막 문장의 '그것'이 뭔지 알려면 앞부분의 '우산'을 기억해야 하는데, 문장이 길어지면 앞부분 정보가 여러 은닉 상태를 반복해서 거치면서 영향이 점점 약해진다. 2. 병렬화 어려움 이전 시점의 은닉 상태가 계산되어야 다음 시점을 처리할 수 있다. 긴 시퀀스에서는 학습 시간이 늘어난다. BPTT RNN은 BPTT(Backpropagation Through Time) 로 학습한다. RNN을 시간축으로 펼친 뒤 오차를 앞 시점으로 전달하는 방법이다. 1. 순전파 — 앞에서부터 예측 계산 2. 손실 계산 — 예측과 정답의 차이 측정 3. 역전파 — 오차를 앞 시점으로 전달하여 가중치 수정 긴 문장에서 생기는 문제는 두 가지다. 학습 신호가 앞부분까지 충분히 전달되지 않을 수 있고, 기울기가 지나치게 작아지거나 커질 수 있다. 2주차 CNN에서 층을 깊게 쌓았을 때 기울기 소실이 생겼던 것과 같은 문제인데, CNN은 공간 방향으로 깊어져서, RNN은 시간 방향으로 길어져서 생긴다는 게 차이다. LSTM과 세 가지 게이트 이 문제를 완화하려고 나온 게 LSTM 이다. 셀 상태와 게이트를 이용해 긴 의존 관계를 더 잘 학습하도록 돕는다. 셀 상태 c_t — 중요한 정보를 비교적 오래 전달하는 경로 은닉 상태 h_t — 현재 출력과 다음 시점에 사용할 정보 게이트 무엇을 결정하는가? 결과 망각 게이트 무엇을 남길까? 이전 정보 유지·삭제 입력 게이트 무엇을 저장할까? 새로운 정보 반영 출력 게이트 무엇을 사용할까? 현재 은닉 상태 생성 게이트는 0과 1 사이 값을 만들어 정보가 통과하는 비율을 조절하는 구조다. 수도꼭지를 여닫는 것처럼 생각하면 될 것 같다. 장점은 기본 RNN보다 긴 의존 관계 학습에 유리하다는 것이고, 한계는 계산 구조가 복잡 하다는 것이다. GRU와 두 가지 게이트 GRU 는 LSTM을 단순화한 게이트 구조다. 별도의 셀 상태 없이 은닉 상태와 두 게이트로 정보의 유지와 갱신을 조절한다. 업데이트 게이트 z_t — 이전 정보와 새 정보를 얼마나 반영할지 결정 리셋 게이트 r_t — 새 상태를 만들 때 과거 정보를 얼마나 반영할지 결정 LSTM보다 구조가 단순하고 같은 은닉 크기에서는 일반적으로 파라미터 수가 적다. 다만 단순하다고 항상 빠르거나 정확한 건 아니고 , 실제 성능은 데이터와 과제에 따라 달라진다고 한다. 전체 흐름 정리 1. 토큰화 이 / 영화는 / 재미있다 2. 토큰 ID 5 / 91 / 13 3. 임베딩 x1 / x2 / x3 4. RNN 계열 순서대로 h(t) 갱신 5. 출력층 긍정 0.93 임베딩 — 토큰을 단어의 특징을 나타내는 작은 숫자 벡터로 변환 RNN — 이전 정보를 전달해 단어의 순서와 문맥을 반영 LSTM과 GRU — 게이트로 필요한 정보를 조절해 장기 의존성 문제를 완화 3. Attention과 Transformer its는 무엇을 가리킬까 The Law will never be perfect, but its application should be just — this is what we are missing, in my opinion. 여기서 its 가 뭘 가리킬까? 답은 Law 다. 우리는 이걸 금방 안다. 그런데 어떻게 알았을까? its 라는 단어만 따로 보면 "그것의"라는 뜻뿐이라 아무 정보가 없다. 문장 전체를 봐야만 알 수 있고, 심지어 뒤에 나오는 단어들까지 봐야 확신할 수 있다. 사람은 이걸 순식간에 하는데 컴퓨터한테는 정말 어려운 문제였다고 한다. 문장이 길어지고 가리키는 대상에서 멀어질수록 앞에 뭐가 있었는지 기억을 못 한다. 멀리 떨어진 단어끼리 관련이 있다는 걸 어떻게 알아낼 수 있을까? 이 질문에서 시작한 게 Attention과 Transformer다. 과거 시퀀스 모델들의 발전 과정 RNN — 순서를 다루게 됐지만 주요 특징 순환 구조로 시간적 순서를 반영하고, 과거의 정보를 현재에 반영할 수 있게 됨 문제점 반복에 따라 과거 정보가 전해지긴 하지만 기울기 소실로 점점 흐려짐 (장기 의존성 문제) 문장을 순차적으로 입력해서 병렬화 불가 LSTM — 기억을 조절했지만 주요 특징 Cell state와 게이트로 "기억할 것"과 "잊을 것"을 조절 RNN보다 오래된 정보를 안정적으로 유지 → 기울기 소실 문제 완화 문제점 게이트 구조 때문에 파라미터가 늘어나 계산이 더 무거움 여전히 단어를 순서대로 하나씩 처리해야 함 → 병렬화 불가능이라는 근본적 한계는 그대로 Seq2Seq — 번역을 가능하게 했지만 I am a person → 저는 사람입니다 처럼 입력과 출력 길이가 다른 작업을 위해 나온 구조다. 인코더가 입력 문장을 읽어 하나의 Context Vector 로 압축하고, 디코더가 그 벡터를 받아 출력 문장을 순서대로 생성한다. 문제점 병렬화 불가능은 그대로 더 치명적인 건, 문장 전체 정보를 벡터 하나에 다 욱여넣어야 해서 문장이 길어지면 초반 정보가 소실 된다는 것 결국 RNN의 한계가 재발한 셈이다. 이걸 해결하려고 나온 게 Attention 이다. RNN : 순서를 다룰 수 있게 됐다 → 멀리 못 보고, 느리다 LSTM : 멀리 볼 수 있게 됐다 → 여전히 순차 처리라 느리다 Seq2Seq : 길이가 달라도 되게 됐다 → 벡터 하나에 압축하는 병목 Attention이 무엇인가 Attention이란 입력 문장의 모든 정보를 균일하게 보는 대신, 필요한 부분만 선택적으로 집중해서 참고하는 방법 --> "필요한 정보에, 필요한 만큼 집중하기" Query와 모든 Key의 유사도를 계산해서, 관련도 높은 Key의 Value를 더 많이 반영한다 문장이 길어도 필요한 정보에만 집중할 수 있다 Query, Key, Value 역할 도서관에서는 Query 어떤 것을 찾고 싶은가? 검색어 Key 나는 무엇을 갖고 있는가? 색인 카드 Value 실제로 무엇을 전달하는가? 실제 내용물 작동 원리는 세 단계다. 1. 주어진 쿼리에 대해 모든 키와의 유사도를 각각 구한다 2. 구한 유사도를 키와 매핑되어 있는 각각의 값에 반영한다 3. 유사도가 반영된 값을 모두 더하여 리턴 ⇒ Attention Value 실제 모델에서는 이렇게 대응된다. Query = t 시점의 디코더 셀에서의 은닉 상태 Key / Value = 모든 시점의 인코더 셀의 은닉 상태들 K와 V를 왜 나누는지 처음엔 몰랐는데, 찾을 때 쓰는 건 제목이고 실제로 읽는 건 본문 이라고 생각하니 조금은.. 이해가... Transformer 구조 개요 원래 Attention은 RNN 위에 얹어서 디코더가 필요한 부분을 더 들여다보게 만드는 보조 장치 였다. 그런데 2017년 구글 논문에서 한 발 더 나아갔다. RNN이나 CNN을 완전히 다 빼버리고 Attention만으로 시퀀스를 처리하면 어떨까? 그렇게 나온 게 Transformer 다. 논문 제목이 "Attention Is All You Need"인 것도 이 뜻이다. RNN/CNN 없이 순환구조와 합성곱 구조를 완전히 배제한, 순수 Attention만으로 구성된 시퀀스 모델 질적으로 우수하면서도 병렬화가 용이하고 훈련시간도 크게 단축 구조 Encoder–Decoder 구조, 각각 N=6개 층 반복 Encoder — 입력 문장을 문맥이 반영된 표현으로 변환 Decoder — Encoder의 출력을 참고해 출력 문장을 순차적으로 생성 Self-Attention 인코더·디코더 내부의 Attention은 Q, K, V가 모두 자기 문장에서 나온다. 그래서 문장 안에서 단어끼리 관계를 파악하고, 이게 RNN의 역할을 대체 한다. 단, 디코더 중간의 Attention만 Q는 디코더, K/V는 인코더 로 다르다. 왜 이 구조가 좋은가 n = 문장 길이, d = 임베딩 차원, k = 커널 크기 Self-Attention은 병렬 연산과 최단 경로를 동시에 확보한다. RNN은 병렬화도 O(n)이고 경로 길이도 O(n)이다. 즉 문장이 길어지는 만큼 처리 시간도 늘고, 첫 단어가 마지막 단어까지 가는 데 그만큼 단계를 거쳐야 한다. Self-Attention은 둘 다 O(1) 이다. 모든 단어를 동시에 계산하고, 어떤 두 단어든 한 번에 연결된다.
Explore drama series hindi content on ELO TV and discover engaging stories through short-form episodes. From interesting characters to dramatic situations, viewers can browse different stories and find content suited to their entertainment preferences. ELO TV makes it convenient to discover Hindi drama stories and watch them directly on your mobile device whenever you have time. 링크텍스트
Cold Work Permit Procedures: Controlling Everyday Workplace Risks Workplace accidents are often associated with dangerous sites, complex machinery, or major plant shutdowns. Yet many incidents can arise from ordinary work. Tasks like loosening connections, checking valves, taking off protective covers, or making routine adjustments may appear low-risk. Once hazards are overlooked, however, even a familiar task can quickly turn into an unsafe event. That is where a Cold Work Permit plays a vital role in workplace safety. It establishes a controlled process for routine activities by recording hazards, required precautions, assigned duties, and authorization steps under a Permit-to-Work (PTW) system. The goal is straightforward: make sure the activity is properly assessed, the necessary safeguards are ready, and the job proceeds under controlled conditions. A Cold Work Permit applies to work that does not involve open flame, sparks, or another ignition source. Because these activities differ from hot work, they generally do not call for dedicated fire watch arrangements or extensive fire prevention measures. However, calling something “cold” should never be interpreted as calling it risk-free. Workers can still face dangers from stored energy, moving equipment, hazardous substances, pressurized systems, or areas where hands or bodies can become trapped or crushed. Typical cold work may include aligning equipment, tightening bolts, calibrating instruments, performing inspections, cleaning, completing Lockout/Tagout (LOTO) activities, and handling routine housekeeping. Whenever the job could create heat or sparks, whether deliberately or by accident, it should instead be classified and controlled as hot work. The value of a Cold Work Permit becomes especially clear when there is no formal permit process in place. In that situation, teams may rely on experience, assumptions, or informal discussions instead of conducting a thorough risk review. The result can be unsuitable PPE, inadequate isolation, or missed information between crews and shifts. Such weaknesses can create unsafe working conditions, disrupt operations, and contribute to failures in meeting workplace safety requirements. An effective Cold Work Permit process brings structure and accountability to everyday activities. It creates a written record of the hazards involved, the measures required to control them, who is responsible for each action, and how long the authorization remains valid. This replaces informal approaches with a consistent process that can be followed, reviewed, and traced, reducing the chance that an important safety measure will be overlooked. At many sites, a cold work permit is valid for a single shift, often covering roughly eight to twelve hours. When the job extends beyond that period, the authorization normally needs to be reviewed and approved again. The renewal process may involve checking the worksite, confirming that safeguards are still effective, and discussing current conditions with the work team. For major maintenance shutdowns, organizations may use campaign permits, but those permits still need periodic validation to keep activities under control. Defined roles are another important part of permit management. The Issuer, sometimes called the Area Authority, makes the work location ready and grants authorization. The Receiver manages execution and checks that the required controls remain in place during the job. Workers must follow the approved precautions and stop work immediately when conditions change, become unsafe, or differ from what was expected. Safety and operations personnel can also perform inspections or audits to confirm that permit conditions are being followed. The cold work permit process normally begins when someone submits a formal request outlining the activity, location, and planned duration. A risk assessment follows, covering potential mechanical, chemical, pressure, ergonomic, and impact-related hazards. Appropriate isolation and LOTO steps are then completed, including locking, tagging, separating energy sources, and verifying that isolation is effective. The worksite is prepared next. This can involve placing barricades, correcting housekeeping issues, and providing adequate lighting. Nearby simultaneous operations (SIMOPS) are reviewed so that one activity does not create a conflict or hazard for another. Personnel receive the required PPE, while tools and equipment are checked for suitability and condition. Before the task starts, the Issuer and Receiver confirm that controls have been established and that workers understand the approved job conditions. During execution, the work environment must continue to be observed for emerging hazards. When an unexpected risk is identified, the activity should stop until the situation has been reviewed and suitable controls are restored. After completion, equipment and systems are returned to their required condition, locks are removed in the proper sequence, and the area is cleaned and checked. The permit can then receive its final authorization and be formally closed. There may not be regulations dedicated solely to cold work permits, but a well-organized permit process can support wider workplace safety obligations. It can help manage areas such as machine guarding, LOTO, hazard communication, PPE, and process safety practices. The permit also creates a written record showing that relevant hazards were considered and controls were established before work began. A useful Cold Work Permit should capture essential details such as the task description, location, equipment, scope, and validity period. It should also identify isolation points, verification steps, barricading needs, guarding measures, housekeeping status, SIMOPS considerations, and any required gas testing. Approval records, restoration steps, and procedures for removing locks should be documented clearly as well. Electronic Permit-to-Work (e-PTW) systems can make this workflow more efficient. Digital platforms can streamline permit creation, use mandatory fields to improve consistency, and automatically capture timestamps for tracking and audits. Centralized dashboards also make active work easier to see, allowing teams to recognize competing activities before they create operational conflicts. This creates a clearer, more reliable, and more efficient approach to permit administration while strengthening the overall management of workplace safety. Book a free demo @ https://toolkitx.com/blogsdetails.aspx?title=Cold-work-permit-(2025-guide)%3A-definition%2C-OSHA%2FHSE-mapping-and-checklist
여러 ai에이전트 활용을 도와주는 Orca를 설치하고 Codex를 연동한 뒤, AGENTS.md 를 작성해서 사용해 보려고 한다. 이번에는 워크스페이스 폴더를 만들고, 내가 어떤 일을 하고 무엇을 알고 있는지 지침 파일에 정리했다. 이후 Codex에 파일을 요약하고 수정하도록 요청하면서 사용해 봤다. 1. Orca 설치 먼저 Orca GitHub 저장소에 들어간 뒤, 다운로드 페이지에서 Windows용 설치 파일을 내려받아 설치했다. Orca는 Codex , Claude 같은 코딩 에이전트를 실행하고 작업을 관리할 수 있는 도구다. 이번에는 Codex를 연결해 온보딩 문서를 작성하는 데 사용했다. 설치 후에는 Orca의 온보딩 체크리스트를 따라 초기 설정을 진행했다. 2. Orca CLI와 Codex CLI 설치 온보딩 과정에서 Orca CLI를 설치했다. CLI는 터미널에서 명령어로 사용하는 인터페이스를 뜻한다. 새 노트북이라 개발 도구가 준비되어 있지 않았고, 기억하기로는 이 과정에서 Node.js와 Visual C++ 관련 구성 요소 를 추가로 설치한 뒤 Orca CLI 설치를 진행했다. 정확한 Visual C++ 패키지명과 설치 순서는 당시 기록이 없어 확실하지 않다. 이후 Codex CLI를 설치하고 로그인 했다. 설치 흐름은 다음과 같다. Orca Windows 앱 설치 ↓ 온보딩 체크리스트에 따라 CLI 설치 준비 ↓ Orca CLI 설치 ↓ Codex CLI 설치 및 로그인 ↓ 온보딩 폴더에서 Codex 사용 이 순서는 내가 진행한 과정을 정리한 것이며, 모든 Windows 환경에서 같은 추가 프로그램이 필요하다는 의미는 아니다. 3. 온보딩 폴더에 AGENTS.md 작성 파일 위치와 역할 문서 폴더 아래에 Codex 온보딩용 디렉터리를 만들고 지침 파일 초안을 작성했다. 당시 대화 기록에서 확인한 위치는 다음과 같다. Documents/ └─ Codex/ └─ 온보딩/ └─ AGENTS.md AGENTS.md 는 Codex에 작업 규칙과 프로젝트 맥락을 전달하는 Markdown 파일이다. Codex는 작업을 시작할 때 적용 대상인 지침 파일을 읽는다. OpenAI 공식 문서 이번 파일에는 코드 작성 규칙뿐 아니라 내 학습 배경과 설명을 받고 싶은 방식 도 넣었다. 처음 작성한 지침 초안의 주요 내용은 다음과 같았다. 아래는 Codex가 당시 파일을 읽고 요약한 내용을 정리한 것이다. 항목 지침 내용 목적 Orca와 Codex의 작업 흐름을 익히는 연습 공간 소통 한국어로 설명하고 개발 용어는 짧게 풀이하기 정보 구분 사실과 추정, 완료한 작업과 미완료 작업 구분하기 파일 수정 기존 파일과 변경 사항을 확인하고 요청한 범위만 수정하기 동작 확인 설치·연동은 실제 실행으로 확인하고, Orca 작업 전 설치 버전의 스킬과 도움말 확인하기 보안 API 키, 토큰, 비밀번호를 문서나 저장소에 기록하지 않기 검증·보고 관련 검사를 실행하고 변경 파일, 확인 방법, 남은 작업을 간결하게 보고하기 4. Codex로 지침을 확인하고 수정하기 파일 요약 요청 처음에는 다음과 같이 요청했다. 현재 위치에 있는 ANGENT.md 파일 요약해줘. 파일명을 잘못 입력했지만, Codex가 현재 폴더에서 관련 이름을 찾아 실제 파일이 AGENTS.md 라는 것을 확인했다. 이후 파일을 읽고 목적, 소통 방식, 수정 원칙 등을 요약했다. 이번에 확인한 기본 지침 파일명은 AGENTS.md 다. 잘못 입력한 이름을 대화 중에 찾아준 것과, 해당 이름의 파일이 자동으로 지침으로 인식되는 것은 별개의 동작이다. 사용자 배경과 학습 방향 추가 초안 확인 후에는 내 업무와 앞으로의 학습 방향을 설명하고 파일을 수정하도록 요청했다. 요청 내용을 정리하면 다음과 같다. AGENTS.md에 내용을 업데이트해줘. 나는 Elastic Stack과 MinIO를 다루는 주니어 데이터 엔지니어야. 앞으로 Elastic Stack을 먼저 배우고 관련 질문을 할 예정이고, 이후에는 MinIO에 관한 질문도 할 수 있어. Codex는 기존 지침을 유지하면서 사용자 배경과 학습 방향 을 추가했다. Elastic Stack과 MinIO를 다루는 주니어 엔지니어 현재는 Elastic Stack을 우선 학습 이후 MinIO로 학습 범위 확장 가능 주니어 눈높이에 맞춰 기초 개념과 업무 예시로 설명 사전 지식과 프로젝트 경험 추가 이어서 기존에 공부했던 내용과 프로젝트 경험도 전달했다. 리눅스 기본 명령어와 Docker, Kubernetes는 기초 정도 알고 있어. 이전에 AWS 클라우드를 공부했고, 온프레미스에서 Kubernetes 환경을 구성해 본 경험이 있어. AWS에서는 EKS를 활용해 MSA를 구성해 봤어. Elastic Stack은 온프레미스 프로젝트에서 ELK를 Helm chart로 간단히 설치하고 사용한 경험이 전부야. 이 내용은 사전 지식과 프로젝트 경험 으로 추가됐다. 구분 전달한 경험 Linux 기본 명령어 사용 Docker·Kubernetes 기초 지식 보유 온프레미스 Kubernetes 환경 구성 AWS 클라우드 학습 및 EKS 기반 MSA 구성 Elastic Stack Helm chart로 ELK를 간단히 설치하고 사용 설명 방식도 함께 보완됐다. Elastic Stack을 설명할 때 기존 AWS·Kubernetes 경험과 연결하되, Elastic Stack의 내부 동작과 운영 개념은 처음부터 설명하도록 반영했다. Codex는 각 수정 후 파일을 다시 읽어 반영된 내용을 확인하고, 무엇을 추가했는지 보고했다. 이번에는 질문에 대한 답변을 받는 데서 끝나는 것이 아니라 실제 Markdown 파일이 수정되는 과정을 확인했다. 주니어 데이터 엔지니어라는 배경과 기존 프로젝트 경험을 지침에 추가했다. 5. 새 세션과 다른 프로젝트에서 활용하기 같은 온보딩 폴더에서 새 세션 시작 같은 폴더를 작업 위치로 선택하면 그 위치의 AGENTS.md 를 활용할 수 있다. 새 세션에서 다음과 같이 요청해 어떤 지침과 배경을 참고하는지 확인할 수 있다. 현재 적용된 지침과 내 학습 배경을 요약해줘. 다른 프로젝트에서 사용 새 프로젝트의 최상위 폴더에 AGENTS.md 를 복사하거나 필요한 내용을 옮긴다. 이미 지침 파일이 있다면 전체를 덮어쓰기보다 사용자 배경, 학습 방향, 사전 지식 처럼 필요한 부분을 기존 내용에 추가하는 방식으로 활용한다. 공통 지침으로 사용 여러 프로젝트에서 공통으로 사용할 내용은 Codex 사용자 설정 폴더인 CODEX_HOME 의 AGENTS.md 에 둘 수 있다. 기본 위치는 ~/.codex/AGENTS.md 이지만, 실행 환경에 따라 달라질 수 있으므로 실제 경로를 확인해야 한다. 전역·프로젝트 지침 안내 당시에는 전역 설정을 변경하지 않고, 현재 파일을 학습용 원본으로 유지한 뒤 다른 프로젝트에 필요한 내용을 가져가는 방식으로 안내받았다. 학습 진도는 별도로 기록 AGENTS.md 에 학습 배경을 적어 두는 것과 이전 대화·학습 진도가 모두 이어지는 것은 다르다. 배운 내용, 막힌 부분, 다음에 할 일은 별도의 학습기록.md 에 정리하는 방법도 안내받았다. AGENTS.md와 학습기록.md를 읽고, 내 수준과 이전 학습 진도에 맞춰 이어서 설명해줘. 이 요청은 앞으로 사용할 수 있는 예시다. 이번 실습에서는 별도의 학습기록 파일까지 만들지는 않았다. 6. 실습 결과 요약 설치 및 연동 Orca 다운로드 페이지에서 Windows용 앱을 설치했다. 온보딩 체크리스트를 따라 Orca CLI를 설치하고, Codex CLI 설치와 로그인을 진행했다. AGENTS.md 작성 및 수정 Documents/Codex/온보딩/AGENTS.md 에 작업 지침을 작성했다. Codex에 파일 요약을 요청해 기존 지침을 확인했다. 주니어 데이터 엔지니어라는 배경과 Elastic Stack 우선 → 이후 MinIO 학습 방향을 추가했다. Linux·Docker·Kubernetes 기초 지식과 온프레미스·EKS·Helm 기반 ELK 사용 경험을 반영했다. Codex가 파일을 수정하고 다시 읽어 변경 내용을 확인하는 과정을 살펴봤다. 재사용 방법 같은 작업 폴더의 새 세션에서는 기존 지침 파일을 활용한다. 다른 프로젝트에는 필요한 지침을 복사하거나 기존 파일에 합친다. 전역 지침 설정과 별도 학습기록 작성 방법은 안내받았으며, 이번에는 실행하지 않았다. 7. 핵심 정리 AGENTS.md 에는 작업 규칙뿐 아니라 사용자 배경과 학습 수준도 정리할 수 있다. 자연어로 수정할 내용을 전달하면 Codex가 실제 파일을 수정하고 결과를 보고할 수 있다. 새 작업에서 사용할 때는 파일명과 위치를 확인하고, 적용 중인 지침을 요약하도록 요청한다. 반복해서 사용할 배경·규칙과 매번 달라지는 학습 진도는 나누어 기록한다.
DQ REVIEW 수정 메서드 : 배열 메서드 중 원본 배열 자체를 직접 변경하는 메서드와 접근자 메서드 : 원본은 그대로 두고 새로운 배열이나 값을 반환하는 메서드를 통칭하는 용어 화살표 함수의 특징 함수 실행 자체 목적, 생성자 역할 불가능, prototype 미존재 클로저 : 내부 함수가 외부 함수의 실행 완료 후에도 소멸된 상위 렉시컬 환경의 변수를 지속적으로 참조하고 조작할 수 있는 메모리 구조 6-1 브라우저 -> 서버에게 요청 서버 -> 브라우저에 응답, html, 텍스트 전달 브라우저 - document 객체 변환 -> 트리 구조 구축 (DOM Tree) [ -> script 태그 -> 동기 실행] -> DOMContentLoaded 이벤트 발생 -> Load 이벤트 발생 ( 자원 로드 완료 ) 웹 브라우저의 자바스크립트 실행 9단계 순서 호스트 객체 런타임 마다 다르게 구성되어 있는 특화 객체 웹브라우저 내에 탑재, 웹브라우저 객체 브라우저의 기능과 밀접하게 관련 window location 객체 : 주소창과 관련된 정보 / 기능을 제공 실시간성, 변경내용이 실제 반영 ( 주소창 변화 ) assign(..) : 지정된 주소로 이동 ( 방문 기록 남김 ) -> location.href -> 속성 변경 시 페이지 이동 replace(..) : 지정된 주소로 이동 ( 방문 기록 남기지 않음 ) reload(..) : 새로고침 history 객체 : 방문기록과 관련된 정보 / 기능 제공 length : 방문 기록 back() : 뒤로 가기 forward() : 앞으로 가기 go() : 음수 = 수 만큼 뒤로 가기, 양수 = 수 만큼 앞으로 가기 scrollRestoration auto : 스크롤 위치 기억 복구 manual : 복구 x scroll 객체 : 브라우저 사용 화면 관련된 정보 / 기능 제공 navigator 객체 : 브라우저 OS 사용환경, 기타 API 관련 정보 / 기능 제공 document 객체 : 문서와 관련된 정보 / 기능 제공 Document 객체 DOM Tree : 빠른 요소 탐색에 필요 요소(DOM)을 선택하는 방법 아이디(id)로 선택 => document.getElementById(아이디명); => 단일 선택 클래스(class)로 선택 => document.getElementsByClassName(클래스명); => 복수 선택 태그(tag)로 선택 => document.getElementsByTagName('태그명'); => 복수 선택 name 속성명으로 선택 => document.getElementsByName('Name 속성명'); => 복수 css 선택자로 선택 => document.querySelector('css 선택자'); => 첫번째 요소 단일 선택 document.querySelectorAll('css 선택자'); => 복수 선택 문서 전체 (html) - window.document head - window.document.head body - window.document.body 양식 (form) 이름 - frmNIDLogin 양식 하위 입력 태그는 이름으로 바로 접근 가능 frmNIDLogin.id 요소 노트 (상대적 접근) parentElement : 부모 요소 children : 자식 요소 firstElementchild : 첫번째 자식 lastElementchild : 마지막 자식 nextElementSibling : 다음 근접 요소 previousElementSiblig : 이전 근접 요소 속성 (Attribute) document.getAttribute('속성명') => 속성 조회 document.setAttribute('속성명', '속성값') => 속성 추가 / 변경 document.removeAttribute('속성명') => 속성 제거 document.hasAttribute('속성명') => 속성의 존재여부 많이 쓰는 속성 바로 접근 가능 id, type, src, href, target, action ... className, style 데이터 속성 다른 속성들의 영향을 끼치지 않고, 데이터를 전달하기 위해 실시간성 보장 document.dataset() 클래스 속성 document.classList 추가, 수정, 삭제, 조회 실습 과제 [6장 1강] 브라우저 렌더링 엔진과 브라우저 객체 모델(BOM) [6장 2강] - DOM 객체 트리 탐색 및 문서 요소 동적 제어
오늘은 펄서 데이터의 KNN 분류를 다시 확인한 뒤, 이진 분류에 사용하는 로지스틱 회귀를 타이타닉과 소셜 광고 데이터에 적용했다. 선형 회귀와 로지스틱 회귀가 같은 0·1 데이터를 다룰 때 왜 다른 선택이 되는지도 비교했다. 펄서 데이터의 불균형 다시 보기 펄서 데이터는 우주 잡음 11,375개와 펄서 1,153개로 클래스 비율 차이가 컸다. 잡음 데이터에서 1,153개만 무작위로 뽑는 언더샘플링으로 두 클래스를 맞춘 뒤, 신호 평균과 표준편차 두 특성으로 KNN을 학습했다. 표준화 후 테스트 정확도는 약 0.892였다. 혼동 행렬에서는 잡음 211건, 펄서 201건을 맞췄고 잡음 20건과 펄서 30건을 반대로 분류했다. 정확도만 보지 않고 어떤 클래스를 틀렸는지 함께 봐야 하는 이유를 확인했다. 타이타닉 전처리와 다중공선성 타이타닉 데이터에서는 ID, 이름, 티켓, 객실 번호, 탑승 항구를 제외하고 Age 결측값을 중앙값으로 채웠다. 성별은 남성 0 , 여성 1 로 변환하고, 동승 가족 수를 합친 FamilySize 컬럼도 만들었다. titanic_df["Age"] = titanic_df["Age"].fillna(titanic_df["Age"].median()) titanic_df["Age"] = titanic_df["Age"].astype(int) titanic_df["Sex"] = titanic_df["Sex"].map({"male": 0, "female": 1}) titanic_df["FamilySize"] = titanic_df["SibSp"] + titanic_df["Parch"] + 1 처음에는 SibSp , Parch , FamilySize 를 모두 입력값으로 넣고 VIF를 계산했다. FamilySize 가 앞의 두 컬럼으로부터 만들어진 값이라 VIF가 매우 크게 나오고 행렬이 rank-deficient라는 경고가 발생했다. titanic_df = titanic_df.drop(["SibSp", "Parch"], axis=1) 원본 두 컬럼을 제거한 뒤에는 다섯 특성의 VIF가 약 1.08~1.68 범위로 내려갔다. 파생변수를 만들 때는 원본 변수와 동시에 넣어 중복된 정보를 만들지 않는지 확인해야 한다. 로지스틱 회귀의 확률 계산 정리한 타이타닉 특성을 8:2로 층화 분리하고 표준화한 뒤 로지스틱 회귀를 학습했다. 학습 정확도는 약 0.799, 테스트 정확도는 약 0.788이었다. lr = LogisticRegression() lr.fit(X_train_scaled, y_train) y_pred = lr.predict(X_test_scaled) proba = lr.predict_proba(X_test_scaled) predict_proba() 로 사망과 생존의 확률을 확인하고, 가중치와 절편으로 선형식 z 를 계산한 뒤 시그모이드 함수를 직접 적용해 봤다. 예시 승객은 생존 확률 약 0.476으로 계산됐고, scikit-learn의 predict_proba() 결과와 같았다. 경사하강법으로 가중치와 절편을 직접 갱신하는 코드도 작성했다. 10,000번 반복 후 테스트 정확도는 라이브러리 모델과 같은 약 0.788이 나왔다. 소셜 광고 데이터와 선형 회귀 비교 소셜 광고 데이터에서는 사용자 ID를 제외하고 성별을 원-핫 인코딩했다. 나이, 추정 연봉, 성별로 구매 여부를 예측한 로지스틱 회귀의 학습 정확도는 약 0.853, 테스트 정확도는 약 0.788이었다. 표준화된 특성의 가중치는 나이 2.39, 연봉 1.16, 성별 0.37 순으로 나왔다. 이번 모델에서는 나이의 가중치가 가장 컸지만, 이는 데이터에서 확인된 모델의 관계이며 구매 원인을 단정하는 결과는 아니다. 같은 구매 여부를 연봉 하나로 선형 회귀에 넣어 보니 학습 점수는 약 0.141이었다. 0·1 범주를 예측하는 문제에는 연속값을 예측하는 선형 회귀보다, 결과를 확률로 변환해 분류하는 로지스틱 회귀가 더 맞는다는 차이를 확인했다. 정리 오늘은 모델을 학습하기 전 변수의 관계를 점검하고, 분류 모델의 확률 계산을 직접 따라가 봤다. 다음에는 ROC 곡선과 혼동 행렬을 함께 보면서 임계값에 따라 분류 결과가 어떻게 달라지는지도 확인해 봐야겠다.