AI Open-Weight Models

  

AI Open-Weight Models

Concept and Modification Pathways

Abstract

Open-weight AI models are trained neural networks whose learned parameters—weights and often biases—are publicly released for download, inspection, and reuse. They occupy a spectrum between closed, API-only systems and fully open-source AI, because the weights may be shared without the original training data, training code, or complete reproducibility pipeline. Their significance lies in enabling local deployment, transparency, reproducibility, domain adaptation, and community-driven innovation. Developers can build on pretrained capabilities through fine-tuning, compression, merging, or architectural extension rather than training from scratch. However, the practical freedom to modify depends on licensing: permissive licenses such as Apache-2.0 or MIT allow broad use, while community or research licenses may restrict commercial deployment or certain applications. Open-weight models—such as Llama, Mistral, Qwen, Gemma, Falcon, and BLOOM—therefore represent both a technical resource and a governance challenge, balancing innovation, auditability, safety, and accountability.

Examples of how open-weight models can be modified:

  • Continued pretraining: Further train on domain corpora, such as legal, medical, financial, or scientific text, to shift general knowledge toward a specialty.
  • Supervised fine-tuning (SFT): Train on instruction–response pairs to improve chat behavior, reasoning style, or task-specific accuracy.
  • Parameter-efficient fine-tuning (PEFT): Use LoRA, QLoRA, adapters, prefix tuning, or IA3 to adapt large models with far fewer trainable parameters.
  • Alignment tuning: Apply RLHF, DPO, ORPO, or KTO to shape helpfulness, harmlessness, honesty, tone, or refusal behavior.
  • Quantization and compression: Convert weights to 8-bit, 4-bit, GPTQ, AWQ, or GGUF formats for cheaper inference; apply pruning, sparsity, or distillation for smaller models.

 

Continued Pretraining (CPT)

Continued Pretraining (CPT) — also called domain-adaptive pretraining or continual pretraining — takes an already pretrained open-weight model and trains it further on a large, mostly unlabeled domain corpus using the same self-supervised objective it was originally trained with. For decoder-only models like Llama, Mistral, or Qwen, that usually means next-token prediction on raw domain text. It is not instruction tuning: there are no question–answer pairs and no task labels. The goal is to shift the model’s internal knowledge, vocabulary, style, and reasoning patterns toward a domain. After CPT, the model is usually instruction-tuned or aligned for actual use.


Legal example

Base model: Llama-3-8B or Mistral-7B (open-weight).

Corpus:

  • Court opinions and case law
  • Statutes, regulations, and codes, e.g. US Code, CFR, EUR-Lex
  • Contracts and SEC exhibits
  • Law reviews, treatises, and legal encyclopedias, where licensing permits
  • Pile of Law or similar open legal datasets

Process:

  1. Clean, deduplicate, and redact PII or privileged material.
  2. Train with causal language modeling on, say, 10–30B legal tokens.
  3. Use a low learning rate, e.g. 1e-5, with warmup and cosine decay.
  4. Mix in 5–10% general web text to reduce catastrophic forgetting.
  5. Optionally extend the tokenizer with legal terms, though this is often unnecessary.

What the model learns:
It becomes better at completing and understanding phrases such as force majeure, indemnification, mens rea, certiorari, consideration, estoppel, and summary judgment. It also absorbs citation formats, contractual structure, and statutory language.

Evaluation:
LegalBench, LexGLUE, ContractNLI, clause extraction, case outcome prediction, and legal summarization. Perplexity on held-out legal text should drop.

Then:
Supervised fine-tuning on legal Q&A, plus retrieval-augmented generation so the model can cite real sources rather than hallucinate cases.

Caution: A CPT legal model is not a lawyer and should not give legal advice. Citations must be verified.


Financial example

Base model: Mistral-7B or Qwen2.5-7B.

Corpus:

  • SEC filings: 10-K, 10-Q, 8-K, S-1, proxy statements
  • Earnings call transcripts
  • Annual reports and investor presentations
  • Financial news and central bank statements
  • Analyst reports, where licensed
  • Accounting standards such as GAAP and IFRS
  • Macroeconomic releases

Process:

  1. Filter for quality, deduplicate, and handle temporal leakage carefully.
  2. Continue pretraining with next-token prediction on 10–50B financial tokens.
  3. Mix general data to preserve broad language ability.
  4. Monitor domain perplexity and downstream financial benchmarks.

What the model learns:
It becomes fluent in EBITDA, accrual, deferred revenue, yield curve, Basel III, derivative hedging, revenue recognition, risk factors, and free cash flow. It also learns the structure of filings and earnings calls.

Evaluation:
FiQA, FinQA, ConvFinQA, financial sentiment, XBRL tagging, earnings call summarization, and numerical reasoning over tables.

Then:
Instruction-tune on financial analysis tasks and connect to tools for live market data, calculators, and filing retrieval.

Caution: Financial CPT models can suffer from outdated knowledge, lookahead bias, and hallucinated numbers. They should not be treated as investment advice.


Combined legal–financial case

SEC filings are a natural overlap. A model continued-pretrained on 10-K risk factors, material contracts, and regulatory exhibits learns both legal and financial language at once. For example, it can better answer:

  • “What litigation risks did the company disclose?”
  • “Summarize the indemnification clause in this exhibit.”
  • “Explain how revenue recognition changed year over year.”

But the same rule applies: CPT gives domain fluency, not professional judgment. It should be paired with retrieval, verification, and human oversight.

 

  • Model merging: Combine multiple fine-tuned variants using SLERP, TIES, DARE, or task arithmetic to blend capabilities.
  • Architecture and context extension: Modify attention, add adapters or experts, expand tokenizer vocabulary for new languages, or use RoPE scaling/YaRN for longer context.
  • Multimodal and tool-use adaptation: Attach vision/audio encoders or train function-calling and agentic behaviors.
  • Knowledge editing and unlearning: Remove specific facts, reduce bias, or erase copyrighted/private information from model behavior.
  • Safety and security modification: Strengthen jailbreak resistance, add watermarking, or red-team and patch vulnerabilities—while noting that the same techniques can also weaken safeguards if misused.

In practice, these modifications are often combined: for example, a base open-weight model can be domain-adapted, LoRA fine-tuned, aligned, quantized, and then merged with another variant for deployment.

Comments