Category report

Malware scanning and binary pattern matching engines

Research date: 2026-10-09.

This selection covers 19 repositories implementing malware signature evaluation, byte-oriented matching, binary classification, process-memory inspection, or substantial file-scanning pipelines. General-purpose matching libraries are included when their byte and streaming semantics directly fit this category. Capability detection and file identification are distinguished from malware verdicts; scanner integrations are included for their own scheduling, extraction, or result-management logic. The descriptions identify useful engineering study material, not uniformly exemplary implementations or independently measured detection effectiveness.

Criteria legend:

  • C1 — Correctness: difficult invariants, concurrency, numerical semantics, adversarial inputs, or failure handling.
  • C2 — Abstractions: substantial reusable interfaces or components serving multiple use cases.
  • C3 — Performance: concrete resource or throughput constraints addressed through understandable architectural choices.
  • C4 — Evolution: sustained development accompanied by compatibility work, tests, or deliberate complexity management; age alone does not qualify.

Canonical repository URLs and archive flags were checked using the GitHub API. Primary source files and documentation were read separately from repository descriptions. None of the selected repositories was marked archived at the research date; explicit maintenance/deprecation notices are identified below. A recent push is not treated as proof of ongoing support. Source links refer to the branches inspected and may change after this date.

Malware signature engines

1. Cisco-Talos/clamav

Language / role: Predominantly C, with C++ and Rust components; antivirus engine, library, command-line scanner, and daemon.

Study libclamav as an example of embedding a hostile-file scanner in different applications. Its API separates engine creation, database loading and compilation, scan options, callbacks, and reference-counted engine lifetime. The byte matcher exposes the additional complexity that real signature languages add to textbook multi-pattern search.

  • C1: The public API explicitly separates safe reference sharing from unsafe mutation of an engine already in use. It also exposes scan-size, file-count, and recursion limits. The matcher validates prefixes, variable-length alternatives, and boundary conditions around candidate matches. These are concrete concurrency and input-boundary obligations, rather than a claim that every parser is safe.
  • C2: File-descriptor and mapped-data scanning, application-defined callback context, and independently configured compiled engines make the scanning subsystem reusable outside the supplied executables.
  • C3: The Aho–Corasick matcher separates candidate finding from forward/backward verification of more complex patterns, making the cost of richer signatures visible in the implementation.

Entry points: Engine and scanning API; Aho–Corasick matcher and pattern verification. The separate fuzzing guide documents scan targets, seed corpora, and dictionaries.

2. VirusTotal/yara

Language / role: C; malware-oriented textual, hexadecimal, and regular-expression rule engine. Status: the repository explicitly declares maintenance mode.

The most instructive subsystem is atom extraction: YARA finds short literal fragments that must occur for a complex pattern to match, scans for those fragments, and then verifies the complete pattern. Its atom-tree explanation connects Boolean coverage of alternatives with the practical selectivity of byte sequences.

  • C1: Choosing atoms must preserve every matching alternative. Selecting just one branch of an alternation could silently lose detections; atoms.c explains why some patterns require multiple atoms and how combinations are represented.
  • C3: Atom quality considers length, byte diversity, and common byte values to reduce expensive verification. This provides a readable example of a correctness-preserving prefilter with workload-sensitive heuristics.
  • C2: The C API separates compiler, compiled rules, and scanner lifetimes; supports memory, files, and file descriptors; and allows application-supplied include resolution and scan callbacks.

Entry points: Atom extraction design and implementation; C API and ownership contracts. Read the maintenance notice before treating it as the locus of new YARA features.

3. VirusTotal/yara-x

Language / role: Rust, with embedding interfaces; a separate implementation of YARA-style malware matching, not a wrapper around libyara.

Study the separation between pattern compilation and rule-condition execution. The compiler maintains pattern atoms and regex code while producing WebAssembly for conditions. Its module interface turns parsed file information into typed values available to rules.

  • C1: Compiler structures account for arbitrary-byte literals separately from UTF-8 source, track ignored-rule dependencies, and deliberately transform global data so compiled rules can be shared across scanning threads. Optional rejection of slow patterns and potentially large loops makes resource policy explicit.
  • C2: The module developer guide specifies Protocol Buffer schemas, Rust entry points, exported functions, and externally packaged modules. This is a reusable extension boundary for different binary formats and analysis domains.
  • C3: Distinct representations for atoms, regex programs, and compiled WebAssembly conditions let each stage optimize its own work rather than interpreting the entire rule uniformly.

Entry points: Compiler representations and build pipeline; Module Developer’s Guide. The project overview distinguishes its development direction from YARA’s maintenance work.

4. vthib/boreal

Language / role: Rust; independent YARA-rule compiler/evaluator with library, CLI, and Python interfaces.

Boreal is useful for studying compatibility implemented through a different internal design. Its matcher selects among complete literal coverage, atom-triggered validation, and direct regex matching. The project documents both its compatibility ambition and deliberate exceptions.

  • C1: ASCII and wide literals have an explicit ordering invariant because inferring the match type from bytes would be incorrect. The compatibility documentation describes comparison against YARA tests and exceptions involving overflow semantics, defensive limits, and self-dependent rules.
  • C3: Matcher construction analyzes regex structure, bypasses atomization for anchored cases, deduplicates literals, and uses a single-match representation to avoid a vector allocation in the common case.
  • C2: Parser, compiler, evaluator, matcher, modules, and process-memory scanning are separate components; feature choices let applications control which format and analysis facilities they include.

Entry points: Matcher selection and literal invariants; Library overview, compatibility exceptions, and scan-elision rationale. Compatibility statements are the project’s documented contract, not an independent equivalence proof.

5. hanul93/kicomav

Language / role: Python; antivirus engine with format-specific plugins, archive handling, signatures, and YARA integration.

This is a useful smaller-community counterpoint to native-code engines. Study the plugin contracts and persistent scan cache, especially how a performance feature becomes part of the scanner’s correctness model.

  • C1: Cache reuse depends on signature version and file metadata, while archive entries distinguish shallow and recursive scanning through a composite key. The implementation also uses thread-local SQLite connections and transactional commit/rollback. These expose concrete invalidation and concurrency questions; metadata-based reuse should not be mistaken for a proof that file contents cannot change.
  • C2: Shared plugin base classes establish initialization, cleanup, rule lookup, scan-size policy, and result conventions across format and detection plugins.
  • C3: Separate file/archive caches, SQLite WAL mode, and cached plugin resources address repeated scanning and concurrent access without hiding the mechanisms behind an external service.

Entry points: Cache schema and invalidation; Plugin interfaces and common behavior. The repository overview establishes the malware-scanning and nested-container scope.

6. intel/hyperscan

Language / role: C/C++; multi-pattern regex matching library for byte buffers and streams, often used in inspection systems.

Study how compilation moves analysis, optimization, and memory sizing out of the scanning path. Immutable databases, mutable stream state, and temporary scratch space have deliberately different lifetimes.

  • C1: Streaming must preserve matches spanning blocks and handle stream completion correctly. The runtime documentation also makes an important concurrency distinction: the library is reentrant, but an individual scratch region cannot be shared by overlapping or recursively invoked scans.
  • C2: One compiled-database model supports block, vectored, and streaming input, with callbacks for matches and serialization for deployment.
  • C3: Scratch and stream-state sizes are determined during compilation, allowing preallocation. The documented execution model avoids repeated allocation during scanning and makes per-stream memory requirements inspectable.

Entry points: Compilation/runtime architecture; Scanning modes, stream handling, and scratch ownership. This entry concerns the verified public GitHub implementation, not an assertion about other Intel distributions.

7. VectorCamp/vectorscan

Language / role: C/C++; a substantively evolved portable fork of Hyperscan.

Counted separately because the project develops architecture abstractions, ARM and Power implementations, SIMD portability, and its own compatibility/testing work. It is particularly useful for examining how vectorized matching algorithms survive changes in vector width, alignment, CPU feature detection, and compiler behavior.

  • C1: The changelog records fixes for chunk-boundary missed matches, SIMD tail masking, out-of-bounds reads, signedness, and runtime CPU detection. These identify real semantic hazards of vectorizing a byte matcher.
  • C3: SIMD dispatch separates common operations from architecture-specific implementations; the project also documents the SuperVector refactoring and alternative SIMDe backend.
  • C4: Dated changes from 2022 through 2026 document portability work, regression tests, build refactoring, and API/ABI commitments. The 5.4.12 entry marks the boundary of full Hyperscan compatibility; later compatibility should not be assumed. The 5.4.13 notes also narrow supported CI platforms.

Entry points: Architecture-specific SIMD dispatch; Independent evolution and compatibility record.

8. strozfriedberg/lightgrep

Language / role: C++ with a C API; binary-oriented multi-pattern regex engine and forensic search CLI. The selected repository contains the library implementation as well as the application.

Lightgrep compiles Unicode-aware patterns into byte patterns for selected encodings rather than requiring the input stream to be valid text. This is valuable for disk slack, memory dumps, mixed encodings, and corrupted files.

  • C1: The API specifies continuity across input buffers, explicit stream closeout/reset behavior, and ordering of hits per keyword rather than globally. Correct consumers must also respect the lifetime relationship between compiled programs and search contexts.
  • C2: Pattern parsing, finite-state-machine construction, immutable bytecode programs, serialization, and independently created search contexts form a reusable compiler/runtime interface.
  • C3: Partial determinization exposes a memory-versus-execution tradeoff, while many contexts can share a compiled program and use bounded context state. The API explicitly warns that serialized programs lack version checks.

Entry points: Library API and lifetime/streaming contracts; Binary and encoding semantics. The older liblightgrep repository is not counted as another implementation.

9. BurntSushi/aho-corasick

Language / role: Rust; reusable multi-pattern literal byte matcher. It belongs here as an engine building block, not as a malware product.

Its unusually substantial design document explains how a compact algorithm develops multiple representations and semantic modes in production. Binwalk’s inspected scanner imports this library, providing a concrete connection to binary identification.

  • C1: Standard, overlapping, leftmost-first, and leftmost-longest matching need different construction/search behavior. Streaming introduces buffering obligations; changing match policy is not merely a presentation choice.
  • C2: A builder and common search interface cover in-memory matching, stream matching, replacements, anchored searches, and ASCII case folding.
  • C3: Noncontiguous NFA, contiguous NFA, and DFA layouts trade allocation, cache locality, and transition-table size. SIMD prefilters and packed literal search provide alternative paths for suitable pattern sets.

Entry points: Internal design and semantic tradeoffs; Public library API overview.

10. CERT-Polska/ursadb

Language / role: C++; n-gram database for querying binary and malware collections.

UrsaDB addresses the other side of scanning performance: avoid rescanning the whole corpus for every rule. Its result is a candidate set requiring exact verification, not a complete YARA verdict.

  • C1: Intersecting per-gram file postings loses ordering information and can yield false positives. The index documentation explicitly presents the prefilter as preserving candidates, while explaining that specialized text indexes cannot answer arbitrary-byte queries on their own.
  • C2: Datasets combine several index types and support tagging and compaction, separating corpus organization from individual query expressions.
  • C3: gram3, text4, wide8, and collision-bearing hash4 indexes trade disk space and indexing work for selectivity. Dataset documentation explains the competing overhead of many small partitions and memory cost of very large ones, especially with wildcards.

Entry points: Index representations and candidate semantics; Dataset architecture and compaction tradeoffs. Individual-file deletion is documented as unimplemented; datasets are the management unit.

Structural signatures, capabilities, and process memory

11. mandiant/capa

Language / role: Python; capability-rule matching over executable features and supported sandbox reports.

Study the distinction between extracting evidence and evaluating rules over it. The matcher operates on features mapped to locations and returns structured result trees, enabling users to inspect why a capability matched. This extends binary-pattern analysis beyond raw byte substrings.

  • C1: Boolean, negation, threshold, and count/range statements must preserve match semantics while returning useful evidence. The engine explicitly distinguishes short-circuited evaluation from complete evaluation and documents when child reordering is safe.
  • C2: A common feature set and statement/result model decouple rule evaluation from analysis backends. The repository overview identifies PE, ELF, .NET, shellcode, and sandbox inputs.
  • C3: Short-circuiting and evaluation counters make rule execution cost visible. The changelog additionally records rule-engine performance work and the expansion of dynamic scopes, alongside compatibility repairs for analysis tools.

Entry points: Rule-evaluation engine; Feature, backend, and matching evolution. Capability evidence is not automatically a maliciousness verdict.

12. hasherezade/pe-sieve

Language / role: C++; Windows process scanner for injected/replaced PE images, shellcode, hooks, and memory patches.

PE-sieve makes an instructive alternative to signatures over disk files. Study how it compares a remote image with an original image while accounting for legitimate loader changes, then groups differences into patches for further analysis.

  • C1: The code scanner relocates the original module to the relevant load base before comparison, clears selected import/export-related differences, handles unequal section sizes and padding, and distinguishes scan errors from suspicious modifications. These steps prevent naive byte comparison from becoming the entire detection policy.
  • C2: A common module-scanner interface and report abstraction support specialized scanners. The code scanner builds patch lists and section-level statuses that later reporting/dumping stages can consume.

Entry points: Code comparison and patch collection; Module-scanner abstraction. The project overview establishes the process-memory and implant-detection scope.

13. JusticeRage/Manalyze

Language / role: C++ engine with YARA rules and auxiliary Python tooling; extensible static PE triage and signature scanning.

Manalyze is useful for studying the boundary between a reusable PE parser and independent detection opinions. Plugins inspect the same parsed object and return threat levels, summaries, and detailed findings rather than owning presentation.

  • C2: Internal plugins and dynamically loaded libraries implement the same analysis interface. The documented external ABI uses creation/destruction functions and an API-version method, allowing plugins with additional dependencies to remain separate.
  • C1: Parser-amplification tests exercise malformed Rich headers and COFF symbol extents, work-budget boundaries, bounded diagnostics, and preservation of later metadata after a limited parse.
  • C3: Those tests check linear Rich-header processing and explicit parser work limits, connecting hostile-input correctness with CPU and diagnostic-volume constraints.

Entry points: Plugin architecture and result contracts; Parser-amplification regression tests. Its overview describes ClamAV-signature, suspicious-import, packer, and cryptographic-constant checks.

14. horsicq/die_script

Language / role: C++/Qt hosting script-based signatures; the Detect It Easy signature-execution subsystem.

This is the implementation-focused entry for the DiE family. Its relationship to the application is verified by DIE-engine’s submodule manifest; the GUI and signature database are not counted separately. Study how format-aware objects are exposed to scripts and how signature execution becomes structured scan results.

  • C2: DiE_ScriptEngine registers utility and format-specific objects and exposes common result, include, logging, and stopping operations. The host supports both QtScript and the alternative QJS-based integration paths.
  • C1: processDetect orders global and format-specific initialization, filters signatures by type and scan options, tracks parent/result identity, propagates script errors into explicit error records, and supports cancellation. These are nontrivial host/interpreter state contracts; this entry does not claim scripts execute in a security sandbox.
  • C3: Signature selection avoids running irrelevant format, deep-scan, or heuristic checks, and optional timing exposes individual signature costs.

Entry points: Signature selection, initialization, and result lifecycle; Script-engine bindings. DiE identifies binary formats and construction characteristics rather than certifying files as benign.

15. ReFirmLabs/binwalk

Language / role: Rust in the inspected v3 implementation; embedded-file and firmware signature scanner/extractor.

Binwalk belongs under binary pattern matching, with firmware analysis as its primary application. Its central lesson is that recognizing a magic-byte sequence is only the start of identifying an embedded object.

  • C1: Signature parsers validate candidate matches and, where possible, determine object extent. The signature interface distinguishes short signatures restricted to the beginning of input and describes rejecting misleading magic matches after structural or checksum validation.
  • C2: Signature attributes, parser functions, result records, and optional extractors form a reusable extension model for many data formats. The library interface returns both identification and extraction results.
  • C3: The scanner uses Aho–Corasick for magic-byte candidate discovery and maps each pattern back to its validating signature, separating broad search from format-specific work.

Entry points: Signature/parser/extractor contract; Primary scanner interface and pattern mapping. This is binary identification, not an antivirus verdict engine.

Scanning systems with substantial surrounding logic

16. target/strelka

Language / role: Python and Go; distributed file-analysis and recursive extraction system with malware-scanning modules.

Study the whole path from intake to typed scanner dispatch, extracted child files, and returned evidence. Its inclusion is for this pipeline implementation, not merely its ability to invoke YARA or ClamAV.

  • C2: Clients, frontends, Redis coordination, and backends have separate roles. MIME, YARA, and inherited file classifications route content to scanners implementing a common interface; extracted objects retain parent and depth information.
  • C1: Backend code distinguishes request, distribution, and scanner timeouts, rejects expired tasks, and imposes extraction-depth limits. Child-file processing preserves provenance while recursively invoking the dispatcher.
  • C3: Distributed workers, a task/file/result coordinator, and optional hash-based result caching address repeated work and cluster throughput. The architecture documentation explains how these pieces can scale independently.

Entry points: Architecture and scanner documentation; Backend, file, and scanner implementation.

17. chainguard-dev/malcontent

Language / role: Go with YARA-X integration; capability/risk scanning and differential analysis of artifacts, archives, and container images.

The distinguishing use case is comparing behavior between releases rather than interpreting every suspicious capability without context. Its surrounding implementation is substantial enough to study independently of the underlying rule engine.

  • C1: Rule-specific scanner pools track borrowers and retire replaced pools only after their last user releases them. Archive code confines writes beneath extraction roots, checks symlink-related traversal, bounds nesting, and detects repeated ancestor content.
  • C3: Compiled-rule reuse and scanner pooling avoid repeated setup. The pool’s fast acquisition path and explicit replacement synchronization make concurrency and reuse costs visible.
  • C2: The same analysis machinery feeds scan, analyze, and differential workflows across files and container/archive contents, as described in the repository overview.

Entry points: Rule cache and scanner-pool lifecycle; Extraction confinement and nested-archive handling. The overview describes differential risk analysis and its contextual assumptions.

18. rfxn/linux-malware-detect

Language / role: Shell with supporting tools; Linux-oriented malware scanner combining hash, hexadecimal, rule-based, and external-engine stages.

This is a useful study of throughput and lifecycle complexity in a shell implementation. The inspected batch engine shares hex extraction between matching stages, separates literal and wildcard processing, and manages workers, checkpoints, and temporary files.

  • C3: Batch hashing, preloaded signature maps, bounded micro-batches, and shared hex buffers reduce process-launch and repeated-extraction costs. The source exposes these mechanisms rather than only asserting faster scans.
  • C1: Chunk-size validation, dedicated file-descriptor ownership, parent-liveness checks, and exit cleanup demonstrate operational invariants. The tests check chunk configurations and temporary-file cleanup.

Entry points: Batch matching workers; Hex-batch regression tests.

Material limitation: The inspected test comments describe a nondeterministic multi-worker hex-scanning issue and force one worker for a case. This is evidence of a difficult failure mode, not evidence that concurrency correctness is fully resolved. Several tests also assert implementation text rather than end-to-end semantics, so their presence alone should not be treated as comprehensive validation.

19. Neo23x0/Loki

Language / role: Python; file/process IOC and YARA scanner. Status: the current README explicitly deprecates this Python implementation and describes inactive maintenance; retained as a historical engineering study.

Loki demonstrates how heterogeneous evidence becomes an explainable triage score: filename patterns, hashes, YARA matches, and process/network observations contribute reasons and severity. Its source also makes the cost and coverage consequences of scan-selection policy visible.

  • C1: The file path handles unreadable files, encoding problems, and YARA exceptions separately from successful matches. Rule metadata is optional, so scoring and reporting provide defaults for third-party rules rather than assuming one schema.
  • C3: Size/type gating determines whether expensive content checks run, while hash IOC lookup uses sorted collections and binary search. This is an explicit performance/coverage tradeoff, not a guarantee that skipped content is clean.

Entry points: Scanner, IOC lookup, scoring, and error paths; Project scope and deprecation notice. Its successor references are not counted as additional researched engines here.

Search coverage and limitations

Discovery used more than six distinct live-web formulations, including malware-engine architecture; independent Rust YARA evaluators; Hyperscan/Vectorscan SIMD matching; binary/forensic streaming regex engines; n-gram malware indexing; PE and process-memory scanners; format-signature and firmware engines; plugin-based antivirus implementations; and distributed/CI artifact scanning. Representative searches included github malware scanning engine YARA ClamAV architecture pattern matching, github binary pattern matching engine Rust YARA boreal yara-x, site:github.com ursadb binary search, site:github.com lightgrep binary regular expression engine, and site:github.com open source antivirus engine signature scanner C KicomAV. Follow-up searches excluded major engines to look for smaller projects; returns increasingly repeated established candidates or produced broader orchestration tools and lightly substantiated new projects.

Verification combined canonical GitHub API responses and repository trees with direct reads of public implementation files, API/design documentation, tests, and changelogs. No candidate code was executed, dependencies installed, or malware samples downloaded. The report’s judgments about what engineers can learn are grounded in those inspected mechanisms; performance and detection claims were not independently benchmarked. Source review was selective rather than a security audit.

Important boundaries and exclusions:

  • Rule-only repositories, bindings, graphical frontends, tutorial antivirus projects, and lists of tools were not retained as separate engines. DiE is represented by its script engine; Lightgrep’s current integrated repository is counted once.
  • Vectorscan is the one explicitly retained fork because its own portability implementation and dated compatibility/testing history justify separate study. Independent YARA rewrites are separate implementations, not libyara bindings.
  • Full sandbox platforms, general reverse-engineering suites, standalone executable parsers, vulnerability scanners, and fuzzy-hash libraries were excluded to keep the focus on scanning and matching engines. UrsaDB, Binwalk, and Aho–Corasick remain clearly labeled adjacent engine components.
  • New antivirus projects surfaced in searches, but feature breadth and self-reported benchmark tables alone did not earn inclusion. The list is a supported selection, not an exhaustive census or a ranking.
  • YARA’s maintenance status, Loki’s deprecation, and the Linux Malware Detect test caveat matter when choosing a project to adopt rather than merely study. C4 is explicitly awarded where inspected dated history supports it; old repositories do not automatically receive that criterion.
Continue exploringBack to the collection →