AI Open-Weight Models
Concept and Modification Pathways
Abstract
Open-weight AI models are trained neural networks whose
learned parameters—weights and often biases—are publicly released for download,
inspection, and reuse. They occupy a spectrum between closed, API-only systems
and fully open-source AI, because the weights may be shared without the
original training data, training code, or complete reproducibility pipeline.
Their significance lies in enabling local deployment, transparency,
reproducibility, domain adaptation, and community-driven innovation. Developers
can build on pretrained capabilities through fine-tuning, compression, merging,
or architectural extension rather than training from scratch. However, the
practical freedom to modify depends on licensing: permissive licenses such as
Apache-2.0 or MIT allow broad use, while community or research licenses may
restrict commercial deployment or certain applications. Open-weight models—such
as Llama, Mistral, Qwen, Gemma, Falcon, and BLOOM—therefore represent both a
technical resource and a governance challenge, balancing innovation,
auditability, safety, and accountability.
Examples of how open-weight models can be modified:
- Continued pretraining: Further train on
domain corpora, such as legal, medical, financial, or scientific text, to
shift general knowledge toward a specialty.
- Supervised
fine-tuning (SFT): Train on instruction–response pairs to improve chat
behavior, reasoning style, or task-specific accuracy.
- Parameter-efficient
fine-tuning (PEFT): Use LoRA, QLoRA, adapters, prefix tuning, or IA3
to adapt large models with far fewer trainable parameters.
- Alignment
tuning: Apply RLHF, DPO, ORPO, or KTO to shape helpfulness,
harmlessness, honesty, tone, or refusal behavior.
- Quantization
and compression: Convert weights to 8-bit, 4-bit, GPTQ, AWQ, or GGUF
formats for cheaper inference; apply pruning, sparsity, or distillation
for smaller models.
Continued Pretraining (CPT)
Continued Pretraining (CPT) — also called domain-adaptive
pretraining or continual pretraining — takes an already pretrained
open-weight model and trains it further on a large, mostly unlabeled
domain corpus using the same self-supervised objective it was originally
trained with. For decoder-only models like Llama, Mistral, or Qwen, that
usually means next-token prediction on raw domain text. It is not
instruction tuning: there are no question–answer pairs and no task labels. The
goal is to shift the model’s internal knowledge, vocabulary, style, and
reasoning patterns toward a domain. After CPT, the model is usually instruction-tuned
or aligned for actual use.
Legal
example
Base model: Llama-3-8B or Mistral-7B (open-weight).
Corpus:
- Court
opinions and case law
- Statutes,
regulations, and codes, e.g. US Code, CFR, EUR-Lex
- Contracts
and SEC exhibits
- Law
reviews, treatises, and legal encyclopedias, where licensing permits
- Pile
of Law or similar open legal datasets
Process:
- Clean,
deduplicate, and redact PII or privileged material.
- Train
with causal language modeling on, say, 10–30B legal tokens.
- Use
a low learning rate, e.g. 1e-5, with warmup and cosine decay.
- Mix
in 5–10% general web text to reduce catastrophic forgetting.
- Optionally
extend the tokenizer with legal terms, though this is often unnecessary.
What
the model learns:
It becomes better at completing and understanding phrases such as force
majeure, indemnification, mens rea, certiorari, consideration,
estoppel, and summary judgment. It also absorbs citation formats,
contractual structure, and statutory language.
Evaluation:
LegalBench, LexGLUE, ContractNLI, clause extraction, case outcome prediction,
and legal summarization. Perplexity on held-out legal text should drop.
Then:
Supervised fine-tuning on legal Q&A, plus retrieval-augmented generation so
the model can cite real sources rather than hallucinate cases.
Caution: A CPT legal model is not a lawyer and should
not give legal advice. Citations must be verified.
Financial
example
Base model: Mistral-7B or Qwen2.5-7B.
Corpus:
- SEC
filings: 10-K, 10-Q, 8-K, S-1, proxy statements
- Earnings
call transcripts
- Annual
reports and investor presentations
- Financial
news and central bank statements
- Analyst
reports, where licensed
- Accounting
standards such as GAAP and IFRS
- Macroeconomic
releases
Process:
- Filter
for quality, deduplicate, and handle temporal leakage carefully.
- Continue
pretraining with next-token prediction on 10–50B financial tokens.
- Mix
general data to preserve broad language ability.
- Monitor
domain perplexity and downstream financial benchmarks.
What
the model learns:
It becomes fluent in EBITDA, accrual, deferred revenue, yield
curve, Basel III, derivative hedging, revenue recognition,
risk factors, and free cash flow. It also learns the structure of
filings and earnings calls.
Evaluation:
FiQA, FinQA, ConvFinQA, financial sentiment, XBRL tagging, earnings call
summarization, and numerical reasoning over tables.
Then:
Instruction-tune on financial analysis tasks and connect to tools for live
market data, calculators, and filing retrieval.
Caution: Financial CPT models can suffer from
outdated knowledge, lookahead bias, and hallucinated numbers. They should not
be treated as investment advice.
Combined legal–financial case
SEC filings are a natural overlap. A model
continued-pretrained on 10-K risk factors, material contracts, and
regulatory exhibits learns both legal and financial language at once. For
example, it can better answer:
- “What
litigation risks did the company disclose?”
- “Summarize
the indemnification clause in this exhibit.”
- “Explain
how revenue recognition changed year over year.”
But the same rule applies: CPT gives domain fluency, not
professional judgment. It should be paired with retrieval, verification, and
human oversight.
- Model
merging: Combine multiple fine-tuned variants using SLERP, TIES, DARE,
or task arithmetic to blend capabilities.
- Architecture
and context extension: Modify attention, add adapters or experts,
expand tokenizer vocabulary for new languages, or use RoPE scaling/YaRN
for longer context.
- Multimodal
and tool-use adaptation: Attach vision/audio encoders or train
function-calling and agentic behaviors.
- Knowledge
editing and unlearning: Remove specific facts, reduce bias, or erase
copyrighted/private information from model behavior.
- Safety
and security modification: Strengthen jailbreak resistance, add
watermarking, or red-team and patch vulnerabilities—while noting that the
same techniques can also weaken safeguards if misused.
In practice, these modifications are often combined: for
example, a base open-weight model can be domain-adapted, LoRA fine-tuned,
aligned, quantized, and then merged with another variant for deployment.
Comments
Post a Comment