ศูนย์ปฏิบัติการความปลอดภัย (SOC) ต้องเผชิญกับปริมาณ alert มหาศาลทุกวัน องค์กรขนาดกลางที่รัน Wazuh บน endpoint หลายร้อยเครื่อง บวกกับ firewall อย่าง FortiGate ที่ส่ง syslog เข้ามา สามารถสร้าง event ได้หลักหมื่นรายการต่อวันโดยง่าย การทำ correlation ด้วย rule แบบดั้งเดิมจับรูปแบบที่รู้จักได้ดี แต่มักพลาด "หางยาว" ของภัยคุกคาม เช่น alert หลายรายการที่แยกกันดูไม่เกี่ยวข้องแต่จริง ๆ เชื่อมโยงกัน การลาดตระเวนแบบช้า ๆ (low-and-slow reconnaissance) หรือบริบทที่นักวิเคราะห์มนุษย์เท่านั้นจะมองออกว่าน่าสงสัย
Using Small Language Models to Detect Cyber Attacks from Wazuh Logs and Firewall Syslog
Security operations centers drown in alerts. A mid-sized enterprise running Wazuh across a few hundred endpoints, plus a FortiGate or similar firewall forwarding syslog, can easily generate tens of thousands of events per day. Traditional rule-based correlation catches known patterns, but it struggles with the long tail: alert bursts that are technically distinct but semantically related, low-and-slow reconnaissance, or context that only a human analyst would recognize as suspicious.
This is where a small language model (SLM) — not a massive frontier model, but a compact 7B–14B parameter model running on-premises — earns its place in the pipeline. It doesn’t replace Wazuh’s detection engine or your SIEM correlation rules. It sits downstream of them, turning structured alert noise into analyst-ready reasoning.
Why an SLM, not an LLM
For SOC workloads, three constraints usually rule out large hosted models:
- Data residency. Enterprise clients in regulated sectors rarely allow raw security logs to leave their network, which rules out most API-based frontier models.
- Latency and cost at volume. Scoring thousands of alert clusters a day needs a model that’s cheap and fast per call, not one billed per token at frontier rates.
- The task doesn’t need frontier reasoning. Classifying an alert burst against MITRE ATT&CK techniques, or writing a two-sentence incident summary, is well within reach of a well-prompted 7B–14B model.
Models like Llama 3.1 8B, Qwen2.5 14B, or Phi-4, served locally through vLLM or Ollama, hit the right balance: good enough instruction-following for structured JSON output, small enough to run on a single GPU inside the client’s own environment.
Where the SLM sits in the pipeline
The mistake to avoid is feeding raw logs directly into the model. Wazuh and your firewall already do normalization — use that. The SLM’s job starts after aggregation, not before it.
flowchart TD
A["Wazuh agents + FortiGate syslog"] --> B["Wazuh manager: rule matching & normalization"]
B --> C["OpenSearch: alert aggregation by source IP / time window"]
C --> D["SLM scoring service: classify, summarize, map to MITRE ATT&CK"]
D --> E["DFIR-IRIS: enriched case creation"]
E --> F["Shuffle SOAR: branch on SLM verdict"]
F --> G["PagerDuty: human analyst escalation"]
Wazuh alerts are already structured JSON — rule ID, severity level, agent, source and destination IPs, and the raw log line. Aggregate correlated alerts into short windows (say, five minutes, grouped by source IP or user) before they ever reach the model. This keeps the SLM reasoning over compact event sequences instead of parsing firehose text, which both lowers cost and improves accuracy.
What the SLM is actually good at
Rule engines are precise but brittle. They’re excellent at "this exact log pattern occurred," and weak at "this sequence of otherwise-unremarkable events looks like an attack in progress." That gap is where an SLM adds value:
- Incident summarization. Turning a cluster of twelve related Wazuh alerts into a two-sentence plain-language summary an on-call analyst can read in five seconds, instead of scrolling through raw logs.
- Pattern reasoning across log sources. Correlating a burst of FortiGate connection denials with a subsequent run of Wazuh authentication failures on the same source IP, and describing why that combination resembles password spraying or credential stuffing — reasoning that static correlation rules often miss unless someone wrote a rule for that exact combination in advance.
- False-positive suppression. Scoring alerts against short embedded context (asset criticality, known maintenance windows, expected admin behavior) to cut noise before a case is even opened in DFIR-IRIS.
Structured output, not free text
The model should never hand back a paragraph for a SOAR platform to parse with regex. Prompt it to return strict JSON:
{
"severity": "high",
"mitre_technique": "T1110.003 - Password Spraying",
"summary": "17 failed logins across 4 accounts from a single external IP within 6 minutes, followed by one successful login.",
"recommended_action": "disable_account_and_escalate",
"confidence": 0.82
}
Shuffle can then branch directly on severity and recommended_action fields, routing high-confidence, high-severity cases straight to PagerDuty while lower-confidence ones queue for analyst review.
Guardrails matter more than accuracy
The single most important design decision is what the model is not allowed to do. An SLM in a SOC pipeline should be advisory only — it summarizes, classifies, and recommends, but it never auto-closes a case or auto-remediates a system. Human review stays in the loop for anything above a defined severity threshold. This isn’t just caution for its own sake; it’s the difference between a defensible SOC process and one that can’t answer for a false negative during a client audit.
Getting started without labeled data
Most organizations don’t have a clean dataset of "alert cluster → correct verdict" pairs to fine-tune on, and that’s fine to start. Strong few-shot prompting — a handful of example alert clusters with their correct MITRE mapping and severity, embedded directly in the system prompt alongside your organization’s specific rule descriptions — gets a 7B–14B model to reasonable accuracy immediately. Fine-tuning becomes worthwhile once you’ve accumulated a few hundred real, analyst-verified incidents to train on.
The result is a SOC pipeline where Wazuh and your firewall still do what they’re best at — fast, deterministic detection — while the SLM handles the messier, more contextual layer of reasoning that used to require a tired analyst squinting at a dashboard at 2 a.m.
如何用 Django 从零搭建 ERP:数据模型、工作流与系统架构
对大多数公司来说,Odoo 和 ERPNext 已经能够很好地解决 ERP 的问题。如果你的业务流程能落在标准模块配置选项的 20% 之内,那么定制现有平台几乎总是比从零构建更划算——我们之前撰写的 ERPNext 实施指南、以及关于 ERP 项目为何失败的文章,正是基于这个原因。
Djangoでゼロから作るERP:データモデル、ワークフロー、アーキテクチャ
OdooやERPNextは、ほとんどの企業にとってERPの課題を十分に解決してくれます。自社の業務プロセスが標準モジュールの設定項目の20%程度に収まるのであれば、既存プラットフォームをカスタマイズするほうが、ゼロから構築するよりもほぼ確実に安上がりです——私たちがERPNextの導入ガイドやERPプロジェクトが失敗する理由について記事を書いてきたのも、まさにこの理由からです。
วิธีสร้างระบบ ERP ตั้งแต่ต้นด้วย Django: โมเดลข้อมูล เวิร์กโฟลว์ และสถาปัตยกรรม
Odoo และ ERPNext แก้ปัญหา ERP ให้กับบริษัทส่วนใหญ่ได้อยู่แล้ว หากกระบวนการทำงานของคุณครอบคลุมอยู่ในตัวเลือกการกำหนดค่าของโมดูลมาตรฐานสัก 20% การปรับแต่งแพลตฟอร์มที่มีอยู่แล้วมักจะถูกกว่าการสร้างใหม่เสมอ — เราเคยเขียนเรื่องการติดตั้งใช้งาน ERPNext และทำไมโครงการ ERP ถึงล้มเหลว ไว้ด้วยเหตุผลนี้เอง
Continue reading “วิธีสร้างระบบ ERP ตั้งแต่ต้นด้วย Django: โมเดลข้อมูล เวิร์กโฟลว์ และสถาปัตยกรรม”
How to Build an ERP From Scratch With Django: Data Model, Workflow, and Architecture
Odoo and ERPNext solve the ERP problem for most companies. If your processes fit within 20% of a standard module’s configuration options, customizing an existing platform is almost always cheaper than building one — we’ve written about implementing ERPNext and about why ERP projects fail for exactly that reason.
如何让 Odoo 或 ERPNext 重新变快:实用性能排查指南
每一套 ERP 系统在上线第一天都很快。六个月后,加上四十个自定义字段,销售团队就开始抱怨 Sales Order 列表视图要加载八秒钟,还有人已经悄悄开始用一份 Excel 表格"临时顶一下,等系统修好再说"。这是我们最常收到的支持工单之一,而且几乎从来不是单一原因造成的——通常是三四个小问题叠加在一起。
Odoo・ERPNextを再び高速化する方法:実践的パフォーマンス改善ガイド
どのERPも導入初日は快適に動きます。しかし6か月後、カスタムフィールドが40個増えた頃には、営業チームからSales Orderの一覧画面が表示されるまで8秒かかると苦情が来て、誰かがひそかに「直るまでの代わり」としてExcelを使い始めている——これはよくあるサポートチケットのひとつであり、原因が一つだけということはほとんどありません。たいていは小さな原因が3つか4つ積み重なっています。
วิธีทำให้ Odoo หรือ ERPNext เร็วขึ้นอีกครั้ง: คู่มือแก้ปัญหาประสิทธิภาพเชิงปฏิบัติ
ERPNext ทุกระบบเร็วในวันแรกที่ใช้งาน ผ่านไปหกเดือนพร้อมฟิลด์กำหนดเองสี่สิบตัว ทีมขายก็เริ่มบ่นว่าหน้ารายการ Sales Order โหลดนานถึงแปดวินาที และมีใครสักคนเริ่มแอบทำสเปรดชีตสำรอง "ไว้ใช้ชั่วคราวจนกว่าจะแก้ปัญหาได้" นี่คือหนึ่งในทิกเก็ตซัพพอร์ตที่พบบ่อยที่สุด และแทบไม่เคยมีสาเหตุเดียว มักเป็นปัญหาเล็กๆ สามสี่อย่างซ้อนทับกัน
How to Make Odoo or ERPNext Faster Again: A Practical Performance Troubleshooting Guide
Every ERP was fast on day one. Six months and forty custom fields later, the sales team is complaining that Sales Order list views take eight seconds to load, and someone has quietly started keeping a spreadsheet on the side "until it’s fixed." This is one of the most common support tickets we get, and it’s almost never one single cause — it’s three or four small ones stacking on top of each other.
