What do you think?


Domain-Specific Small Language Models: Efficient AI for local deployment
Bigger isn’t always better. Train and tune highly focused language models optimized for domain specific tasks.
When you need a language model to respond accurately and quickly about a specific field of knowledge, the sprawling capacity of a LLM may hurt more than it helps. Domain-Specific Small Language Models teaches you to build generative AI models optimized for specific fields.
In Domain-Specific Small Language Models you’ll
• Model sizing best practices
• Open source libraries, frameworks, utilities and runtimes
• Fine-tuning techniques for custom datasets
• Hugging Face’s libraries for SLMs
• Running SLMs on commodity hardware
• Model optimization or quantization
Perfect for cost- or hardware-constrained environments, Small Language Models (SLMs) train on domain specific data for high-quality results in specific tasks. In Domain-Specific Small Language Models you’ll develop SLMs that can generate everything from Python code to protein structures and antibody sequences—all on commodity hardware.
About the book
Domain-Specific Small Language Models teaches you how to create language models that deliver the power of LLMs for specific areas of knowledge. You’ll learn to minimize the computational horsepower your models require, while keeping high–quality performance times and output. You’ll appreciate the clear explanations of complex technical concepts alongside working code samples you can run and replicate on your laptop. Plus, you’ll learn to develop and deliver RAG systems and AI agents that rely solely on SLMs, and without the costs of foundation model access.
About the reader
For machine learning engineers familiar with Python.
About the author
Guglielmo Iozzia is a Director, ML/AI and Applied Mathematics at MSD. He studied Electronic and Biomedical Engineering at the University of Bologna, has an extensive background in Software and ML/AI Engineering applied to real-life use cases across different industries, such as Biotech Manufacturing, Healthcare, Cloud Operations, and Cyber Security.
When you need a language model to respond accurately and quickly about a specific field of knowledge, the sprawling capacity of a LLM may hurt more than it helps. Domain-Specific Small Language Models teaches you to build generative AI models optimized for specific fields.
In Domain-Specific Small Language Models you’ll
• Model sizing best practices
• Open source libraries, frameworks, utilities and runtimes
• Fine-tuning techniques for custom datasets
• Hugging Face’s libraries for SLMs
• Running SLMs on commodity hardware
• Model optimization or quantization
Perfect for cost- or hardware-constrained environments, Small Language Models (SLMs) train on domain specific data for high-quality results in specific tasks. In Domain-Specific Small Language Models you’ll develop SLMs that can generate everything from Python code to protein structures and antibody sequences—all on commodity hardware.
About the book
Domain-Specific Small Language Models teaches you how to create language models that deliver the power of LLMs for specific areas of knowledge. You’ll learn to minimize the computational horsepower your models require, while keeping high–quality performance times and output. You’ll appreciate the clear explanations of complex technical concepts alongside working code samples you can run and replicate on your laptop. Plus, you’ll learn to develop and deliver RAG systems and AI agents that rely solely on SLMs, and without the costs of foundation model access.
About the reader
For machine learning engineers familiar with Python.
About the author
Guglielmo Iozzia is a Director, ML/AI and Applied Mathematics at MSD. He studied Electronic and Biomedical Engineering at the University of Bologna, has an extensive background in Software and ML/AI Engineering applied to real-life use cases across different industries, such as Biotech Manufacturing, Healthcare, Cloud Operations, and Cyber Security.
376 pages, Paperback
Published May 26, 2026
Ratings & Reviews
Friends & Following
Create a free account to discover what your friends think of this book!
Community Reviews
Displaying 1 - 11 of 11 reviews
July 20, 2026
I recently picked up book reviewing again after a brief hiatus and was initially worried whether my attention span would hold up. Domain-Specific Small Language Models by Guglielmo Iozzia was the perfect pick to get me back into the rhythm. It is a highly engaging, reader-friendly, and crisp guide that had me genuinely looking forward to turning every page.
If you are feeling overwhelmed by the constant hype surrounding massive, closed-source models, this book is the perfect reality check. It provides a brilliant, structured masterclass on how tailored Small Language Models under 10 billion parameters can deliver massive business value cleanly, securely, and cost-effectively. The book maps out a clear path through the technical wilderness, divided neatly into four logical parts covering foundational elements, core mechanics, real-world use cases, and advanced deployment strategies.
A few engineering milestones that stood out to me:
Hands-on Fine-Tuning: The book uses a code-first approach to data preparation via Hugging Face. Chapter 3 is a great example, showing how to fine-tune a humble GPT-2 small model to generate programmatic Python animation code via the Manim engine using a highly curated dataset.
Production Inference Realities: It cuts through cloud compute myths by breaking down the math of memory-bound workloads, showing you how to budget for GPU limits while maximizing speeds using KV caching and DeepSpeed.
Local and Secure Deployment: As someone who loves practical implementation, I loved the open-source references. Setting up quantized GGUF models locally using Ollama, LM Studio, and Jan opens up incredible possibilities for running private, air-gapped workflows right on your laptop without burning cloud credits.
Agentic AI and GraphRAG: The final chapters connect these small models to advanced modern architectures. It details how to build a framework-less RAG pipeline from scratch, implement a fully local GraphRAG pipeline using Mistral 7B and NetworkX, and experiment with Test-Time Compute and GRPO reinforcement learning.
A personal highlight was using Manning's built-in LiveBook AI assistant to quiz myself after the chapters just to see if my memory and accuracy were optimal, pun absolutely intended.
Whether you are a data engineer, machine learning practitioner, or tech leader looking for a clear, actionable guide to building optimized, domain-specific AI systems, you should definitely pick this up. It is a fantastic, timely resource I will be going back to refer to time and again.
If you are feeling overwhelmed by the constant hype surrounding massive, closed-source models, this book is the perfect reality check. It provides a brilliant, structured masterclass on how tailored Small Language Models under 10 billion parameters can deliver massive business value cleanly, securely, and cost-effectively. The book maps out a clear path through the technical wilderness, divided neatly into four logical parts covering foundational elements, core mechanics, real-world use cases, and advanced deployment strategies.
A few engineering milestones that stood out to me:
Hands-on Fine-Tuning: The book uses a code-first approach to data preparation via Hugging Face. Chapter 3 is a great example, showing how to fine-tune a humble GPT-2 small model to generate programmatic Python animation code via the Manim engine using a highly curated dataset.
Production Inference Realities: It cuts through cloud compute myths by breaking down the math of memory-bound workloads, showing you how to budget for GPU limits while maximizing speeds using KV caching and DeepSpeed.
Local and Secure Deployment: As someone who loves practical implementation, I loved the open-source references. Setting up quantized GGUF models locally using Ollama, LM Studio, and Jan opens up incredible possibilities for running private, air-gapped workflows right on your laptop without burning cloud credits.
Agentic AI and GraphRAG: The final chapters connect these small models to advanced modern architectures. It details how to build a framework-less RAG pipeline from scratch, implement a fully local GraphRAG pipeline using Mistral 7B and NetworkX, and experiment with Test-Time Compute and GRPO reinforcement learning.
A personal highlight was using Manning's built-in LiveBook AI assistant to quiz myself after the chapters just to see if my memory and accuracy were optimal, pun absolutely intended.
Whether you are a data engineer, machine learning practitioner, or tech leader looking for a clear, actionable guide to building optimized, domain-specific AI systems, you should definitely pick this up. It is a fantastic, timely resource I will be going back to refer to time and again.
June 29, 2026
I like Manning books, and after reading LLMs in Production and Knowledge Graphs and LLMs in Action I realized that there are gaps in what I knew after reading this book. Nearly the first half of the book is talking about the basics, and then the last half is advanced concepts, which was interesting.
Some of the basics is about quantizing LLMs and fine-tuning LLMs with lots of different ideas of how to fine-tune, and that is great as an intro but then there are libraries and papers I didn't know about, such as Smoothquant, which I expect can also be used for yolo type image recognition and video classification, to make these models work well with Android and other smaller devices so I may be able to get some of my models to run on smaller devices.
I forgot onnx was also in the intro section and now I am wondering if I can use smoothquant on an onnx library, but I expect I need to do smoothquant and then onnx, but it gives me ideas that marks when a book is really good, how can I move beyond what it talked about.
It also talks about Bitnet, but that seems LLM specific and FlexGen which is NVIDIA only, so of limited use to me, right now.
Then they go into rag, graph rag, which is interesting for a domain specific LLM, as this is for having fine-tuned an LLM for your specific needs, and I love the examples of generating a crystals from prompt and creating antibodies, and I love that they show how to validate the output.
There is more, including deploying, creating personal agents. There is so much more but it would require more posts to go into detail, or read the book. :D
Some of the basics is about quantizing LLMs and fine-tuning LLMs with lots of different ideas of how to fine-tune, and that is great as an intro but then there are libraries and papers I didn't know about, such as Smoothquant, which I expect can also be used for yolo type image recognition and video classification, to make these models work well with Android and other smaller devices so I may be able to get some of my models to run on smaller devices.
I forgot onnx was also in the intro section and now I am wondering if I can use smoothquant on an onnx library, but I expect I need to do smoothquant and then onnx, but it gives me ideas that marks when a book is really good, how can I move beyond what it talked about.
It also talks about Bitnet, but that seems LLM specific and FlexGen which is NVIDIA only, so of limited use to me, right now.
Then they go into rag, graph rag, which is interesting for a domain specific LLM, as this is for having fine-tuned an LLM for your specific needs, and I love the examples of generating a crystals from prompt and creating antibodies, and I love that they show how to validate the output.
There is more, including deploying, creating personal agents. There is so much more but it would require more posts to go into detail, or read the book. :D
June 15, 2026
I picked up Domain-Specific Small Language Models because most of the AI discussion today is centered around larger and larger foundation models, while many enterprise use cases actually benefit from smaller, specialized models.
What I liked about this book is that it stays practical. The author covers the entire journey from data preparation and fine-tuning to quantization, optimization, deployment, RAG, and agentic AI patterns. The chapters on ONNX, quantization, and model optimization stood out to me because these are the topics that often determine whether a model can realistically be deployed in production.
In my own work, I've seen domain-specific models deliver strong results for document processing, knowledge retrieval, and other specialized workflows where latency, cost, or privacy requirements make large models less attractive. The book does a good job explaining where SLMs fit and, just as importantly, where they don't. It avoids the common narrative that SLMs will replace LLMs. In reality, both have their place.
If I had one suggestion, it would be to include more enterprise case studies outside the scientific and code-generation examples. Seeing additional examples from industries such as legal, healthcare, financial services, or cybersecurity would make the concepts even more relatable for a broader audience.
Overall, I found this to be a valuable read for AI practitioners and architects who want to understand not just how to build domain-specific SLMs, but how to optimize and deploy them effectively in real-world environments.
What I liked about this book is that it stays practical. The author covers the entire journey from data preparation and fine-tuning to quantization, optimization, deployment, RAG, and agentic AI patterns. The chapters on ONNX, quantization, and model optimization stood out to me because these are the topics that often determine whether a model can realistically be deployed in production.
In my own work, I've seen domain-specific models deliver strong results for document processing, knowledge retrieval, and other specialized workflows where latency, cost, or privacy requirements make large models less attractive. The book does a good job explaining where SLMs fit and, just as importantly, where they don't. It avoids the common narrative that SLMs will replace LLMs. In reality, both have their place.
If I had one suggestion, it would be to include more enterprise case studies outside the scientific and code-generation examples. Seeing additional examples from industries such as legal, healthcare, financial services, or cybersecurity would make the concepts even more relatable for a broader audience.
Overall, I found this to be a valuable read for AI practitioners and architects who want to understand not just how to build domain-specific SLMs, but how to optimize and deploy them effectively in real-world environments.
May 29, 2026
If you are an AI developer / Researcher or GenAI/SLM enthusiast and struggling with questions like -
1. How to align generalists open-source models to domain specialist knowledge or datasets?
2. How to reduce LM inference cost efficiently using bleeding-edge technology like vLLM, Ollama ?
3 . How to deploy and serve Small Language Models securely and efficiently on commodity hardware ?
Then, the book Domain-Specific Small Language Models by Guglielmo Iozzia published by Manning publications is the one-stop reference for all your queries.
As a reviewer, I can attest from personal experience that the author has put together an amazing amount of information in this book covering some of the hard to follow topics in an easy to understand and practical manner, flattening the learning curve for learners.
Some of the advanced techniques for efficiently executing LLMs on commodity hardware or laptop were explained comprehensively.
The beauty of this book is that it explores the integration of SLM into RAG systems and agentic workflows in a very attractive and hands-on manner.
To conclude, if you desire to train custom language models for specific domain having small-footprint then, this is a must resource on your book-shelf.
1. How to align generalists open-source models to domain specialist knowledge or datasets?
2. How to reduce LM inference cost efficiently using bleeding-edge technology like vLLM, Ollama ?
3 . How to deploy and serve Small Language Models securely and efficiently on commodity hardware ?
Then, the book Domain-Specific Small Language Models by Guglielmo Iozzia published by Manning publications is the one-stop reference for all your queries.
As a reviewer, I can attest from personal experience that the author has put together an amazing amount of information in this book covering some of the hard to follow topics in an easy to understand and practical manner, flattening the learning curve for learners.
Some of the advanced techniques for efficiently executing LLMs on commodity hardware or laptop were explained comprehensively.
The beauty of this book is that it explores the integration of SLM into RAG systems and agentic workflows in a very attractive and hands-on manner.
To conclude, if you desire to train custom language models for specific domain having small-footprint then, this is a must resource on your book-shelf.
June 30, 2026
Rather than rent time on someone else’s large language model, you can use the techniques in the book to
- Download open-weight LLMs
- Customize them to your specific tasks
- Integrate them into your computing environment
- Run them on your end devices and servers.
Running customized models on your own infrastructure may be less expensive and, because you control the model, enables you to determine the model’s effectiveness without having to cope with updates provided by a hosted model.
The book walks through many case studies, providing working code for each.
The tools and techniques covered include
- LoRA, to fine-tune the weights of an already trained model
- ONNX, a file format for model weights and computation, and a tool set for that format. ONNX can quantize models, which reduces their memory and computation footprints, often reducing costs and latency with acceptable accuracy
- Ollama, a platform for running models locally
- LM Studio, a Python API for accessing models from your application
- GraphRAG, Microsoft’s implementation of a knowledge graph for RAG queries.
This book is a great resource for builders of locally run, customized applications.
Me: roylowrance.com
- Download open-weight LLMs
- Customize them to your specific tasks
- Integrate them into your computing environment
- Run them on your end devices and servers.
Running customized models on your own infrastructure may be less expensive and, because you control the model, enables you to determine the model’s effectiveness without having to cope with updates provided by a hosted model.
The book walks through many case studies, providing working code for each.
The tools and techniques covered include
- LoRA, to fine-tune the weights of an already trained model
- ONNX, a file format for model weights and computation, and a tool set for that format. ONNX can quantize models, which reduces their memory and computation footprints, often reducing costs and latency with acceptable accuracy
- Ollama, a platform for running models locally
- LM Studio, a Python API for accessing models from your application
- GraphRAG, Microsoft’s implementation of a knowledge graph for RAG queries.
This book is a great resource for builders of locally run, customized applications.
Me: roylowrance.com
May 29, 2026
I was fortunate to get an early opportunity to review this book while the author was still in the process of writing it - and I am glad I did.
If you want to learn how to fine-tune a language model not just on random data but on a domain-specific dataset, and do it on cost-effective hardware, this is the book that technically teaches you exactly how. It covers a great range of topics - from fine-tuning techniques like LoRA and quantization, all the way to serving your model using frameworks like ONNX and vLLM.
I will be honest - this is not a casual read. It is a tech-heavy book that demands your full attention. But that's also what makes it valuable. Whether you are an academic, a hobbyist experimenting on your own machine, or someone building domain-specific models for enterprise use - there is something genuinely useful in here for you.
I definitely recommend this book to anyone serious about building and deploying efficient, domain-specific language models in the real world.
If you want to learn how to fine-tune a language model not just on random data but on a domain-specific dataset, and do it on cost-effective hardware, this is the book that technically teaches you exactly how. It covers a great range of topics - from fine-tuning techniques like LoRA and quantization, all the way to serving your model using frameworks like ONNX and vLLM.
I will be honest - this is not a casual read. It is a tech-heavy book that demands your full attention. But that's also what makes it valuable. Whether you are an academic, a hobbyist experimenting on your own machine, or someone building domain-specific models for enterprise use - there is something genuinely useful in here for you.
I definitely recommend this book to anyone serious about building and deploying efficient, domain-specific language models in the real world.
May 28, 2026
Domain Specific LLMs in Action fills a real gap in the LLM book market by focusing on what most titles skip: how to actually run small, specialized language models well on commodity hardware. The author moves confidently across the full inference stack — ONNX, quantization, profiling, and modern serving with vLLM, FastAPI, and MLC LLM (including a refreshingly practical Android deployment chapter) — with hands-on PyTorch and HuggingFace examples throughout. The two domain showcases, Python code generation and chemistry/protein structures, are an inspired pairing that proves the techniques generalize well beyond a single use case. If you're a practitioner trying to put a tuned, domain-specific LLM into production without a hyperscaler-sized budget, this is one of the most useful books you can pick up
June 4, 2026
I should start by saying that this book isn’t for everyone, and it may not be immediately easy to digest.
The author makes an effort to dive deep into all the subjects that are the foundation of LLMs, but of course it is not possible to cover everything. So, an integration while reading could be needed.
That said, the book is USEFUL.
At a time like this, when we are all spending heavily in AI, we are increasingly building a dependency on rules and usage of models shaped by a handful of major global players.
This book helps you design sharper models tailored to your own domain, and in particular models that are sustainable and more accessible to a wider audience, not just big players.
The knowledge it offers is definitely a valuable card to have in your deck.
The author makes an effort to dive deep into all the subjects that are the foundation of LLMs, but of course it is not possible to cover everything. So, an integration while reading could be needed.
That said, the book is USEFUL.
At a time like this, when we are all spending heavily in AI, we are increasingly building a dependency on rules and usage of models shaped by a handful of major global players.
This book helps you design sharper models tailored to your own domain, and in particular models that are sustainable and more accessible to a wider audience, not just big players.
The knowledge it offers is definitely a valuable card to have in your deck.
April 11, 2026
It took me forever to get through this book. Mostly because I can only read it when I have quiet time and am focused. Iow, not before bed, so typically weekend mornings. And I still need to go through the hands on examples.
The content is excellent. The book illustrates how to explore using and fine tuning LLMs with limited resources with open source tools. The author finds the right level detail in explaining fine tuning and optimization techniques without diving too deep into AI/LLM theory. Chapters typically introduce a concept and the walk through an example.
The content is excellent. The book illustrates how to explore using and fine tuning LLMs with limited resources with open source tools. The author finds the right level detail in explaining fine tuning and optimization techniques without diving too deep into AI/LLM theory. Chapters typically introduce a concept and the walk through an example.
Review of advance copy received from Publisher
A great introduction in how to make your LLMs so small, that you can run them locally. Could have more content about the domain-specific part, nevertheless, a helpful book.
December 3, 2025
great read in understanding small language model for limited computing power and resources.
Displaying 1 - 11 of 11 reviews



