Prompt optimisation
Angle: Declarative prompt specification vs. programmatic prompt composition
promptel is a specification language and tooling; DSPy is a Python framework for compiling prompts. promptel optimises for portability, version control, and human review; DSPy optimises for runtime bootstrapping. They are complementary, not competitors — promptel can express a DSPy signature, and DSPy can compile a promptel prompt into a teleprompter.
Agent protocols
Angle: Compliance + audit layer on top of MCP and A2A
mpl is not a replacement for MCP or A2A; it is a layer that sits between your agents and the underlying transport. Where MCP and A2A define how agents talk, mpl defines what correct looks like — versioned semantic-type contracts, quality scores, BLAKE3-hashed audit records, and policy enforcement. Those records map to common regulatory asks (SOX, GDPR, HIPAA, the EU AI Act), but they are evidence primitives, not certifications: your compliance programme still owns the mapping.
LLM routing
Angle: Self-hosted gateway with per-prompt analytics and trace-driven prompt optimisation
LiteLLM, Portkey, and OpenRouter are mature gateways with broad provider coverage. route-switch is a smaller self-hosted Go gateway: it routes requests across providers by configured strategy, logs success, cost, and latency per prompt to DuckDB, and reruns MIPROv2 prompt optimisation against captured production traces. It publishes no benchmark numbers; the point is that you can measure the cost–quality trade-off on your own traffic.
Sandboxing
Angle: Selective-denial container runtime vs. userspace kernels and microVMs
gVisor is a mature userspace kernel that emulates syscalls; Firecracker is a microVM with a KVM guest; WASM is portable but constrained. zviz is none of these: it is an OCI-compatible Zig runtime that layers namespaces, dropped capabilities, Landlock, seccomp-BPF, and cgroups v2, lets 132 syscalls reach the host kernel natively, and denies 24 dangerous ones. That trades gVisor's deeper isolation for native-speed allowed syscalls and a single static binary; if a workload needs ptrace, mount, or unshare inside the container, use gVisor.
Vector search
Angle: Embedded SQLite-backed ANN vs. dedicated vector databases
For sub-100K-vector corpora and prototype workloads, memista runs as a small local HTTP service, or mounted inside your own Actix app, with no separate database to operate. It is experimental and tested below about 100k vectors. For million-vector corpora with high QPS, Pinecone, Qdrant, or Milvus are better choices. memista is the right answer to "do I really need a vector database?"
Skelf
numaperf
vs Alternative
hwloc, libnuma, Linux sched_setaffinity, custom schedulers
NUMA-aware scheduling
Angle: NUMA-first Rust runtime for latency-critical services vs. general-purpose tools
hwloc and libnuma are the general-purpose NUMA libraries. numaperf is a Rust runtime built on the same ideas for latency-critical services: topology discovery, RAII thread pinning, explicit placement policies, per-node scheduling and sharded structures, device locality, and locality observability. It publishes no benchmark of its own; its observability primitives exist so you can measure whether NUMA placement moves your p99 on your workload.
NL to constraint satisfaction
Angle: LLM-to-formal-solver pipeline vs. pure LLM output
Pure LLM output cannot guarantee that a schedule or assignment is even feasible. savanty uses a DSPy-orchestrated LLM to translate an English description of a discrete constraint problem into Answer Set Programming, then hands it to the Clingo solver, which either finds a valid answer set or proves none exists. The guarantee covers the translated program, not the translation itself; a typed repair loop handles syntax errors, unsatisfiable programs, and empty results. If you can already model in CP-SAT, you may not need it.
Ranking with sparse feedback
Angle: Adaptive pair selection with Elo ratings vs. exhaustive or random pairing
Bradley-Terry and TrueSkill are rating models; the expensive part of a human or LLM-judge ranking exercise is usually choosing which pairs to ask about. compere is a Python package and FastAPI service that uses UCB1 to choose the next pair adaptively and Elo to update ratings. It does not implement Bradley-Terry or TrueSkill, and it publishes no claim about how many comparisons it saves on your data; that depends on the items and the judges.
Skelf
slorg
vs Alternative
Algolia, Meilisearch, Typesense, Elasticsearch + LLM
Deliberative search
Angle: Reasoning before retrieval vs. retrieval-then-ranking
Traditional search retrieves documents that match the query and ranks them. slorg runs a fixed six-step pipeline: an LLM drafts an answer, a knowledge graph is extracted from the draft, search keywords are derived from the graph, SearxNG retrieves candidates, pages are fetched, and each result is scored 0–1 against the original query. It is not an agent and publishes no precision benchmark; the trade is visible intermediate artefacts at the cost of three LLM round-trips per query.
Browser-extension LLM frameworks
Angle: Open framework for LLM browser extensions vs. vendor SDKs
Most LLM browser extensions lock you to a single provider. anouk is a portable framework that abstracts the manifest v3 quirks and provides a unified API for the LLM call — you can swap OpenAI, Anthropic, or local llama.cpp without rewriting the extension.
Ephemeral credentials
Angle: Ephemeral credential proxy for LLM APIs vs. general-purpose key management
General-purpose API gateways and secret managers protect keys on servers you control; they do not help when a browser app needs to call an LLM directly. perishable is a self-hosted Node proxy plus browser SDK: the upstream key stays server-side, and the browser gets short-lived JWT sessions tied to a device fingerprint, with per-fingerprint rate limiting. It raises the cost of casual abuse, not of a determined attacker, and it is not an observability or audit layer.
Skelf
waremax
vs Alternative
RAWSim-O, ARENA-Sim, gym-dispatch, custom DES
Warehouse robotics simulation
Angle: Deterministic RL-ready RMFS simulator vs. legacy simulators
RAWSim-O and ARENA-Sim are the legacy RMFS simulators. waremax is built around three properties they do not jointly provide: exact determinism (byte-identical replay), a first-class RL interface (Gymnasium + PyO3), and instrumented delay attribution usable as a reward signal.
Text-to-video pipelines
Angle: An open, resumable multi-model pipeline vs. a hosted generative video service
Runway, Pika and Stable Video Diffusion synthesise motion and will produce a better-looking clip with almost no setup; direktor does not synthesise video at all, it composes narrated stills. What it offers instead is the pipeline: six separately resumable stages across script generation, BARK narration, Distil-Whisper transcription for timing recovery, FLUX stills and FFmpeg composition, every one of which can be swapped or inspected. Pick direktor when the orchestration is the thing you want to own, and a hosted service when the output is.
Metadata-private messaging
Angle: Authenticated metadata privacy vs. content encryption and transport anonymity
Signal protects content extremely well and assumes the server is trusted not to abuse routing metadata, which means it still sees the social graph. Tor anonymises the network path of a stream and says nothing about who the other party is. tessera authenticates the sender with a Schnorr proof over a per-recipient blinded pseudonym while giving a network observer only differentially-private bucket counts — a narrower claim than "private messaging", and it does not protect content on its own or survive a compromised endpoint.
Skelf
Skelf Research (overall)
vs Alternative
Big Tech AI labs, AI startups, independent researchers
Independent AI research labs
Angle: Independent research lab publishing inspectable, tested software with stated maturity
Big Tech AI labs publish papers; AI startups publish products; Skelf Research publishes runnable, testable code with its maturity, scope, and limitations stated. The methodology — "hypotheses as software" — is the differentiator. We do not publish benchmark numbers we have not measured; every quantitative claim we make is listed with its source in our evidence register.