Our Consultants' Former Clients

Program
7 modules
Format
On-site, online or hybrid
Prerequisites
Python, Linux and Docker
Certificate
Attendance certificate; Boğaziçi University-endorsed on request

Cloud API or Local LLM?

Either can be the right answer. The decision depends on how sensitive the data is, the volume of use and your team's operational capacity. The training teaches the criteria for making it.

Criterion Cloud API Local LLM
Data Requests are processed on the provider's servers. Questions, documents and answers stay inside your network.
Cost You pay per token; the bill grows with usage. You invest in hardware up front; capacity and cost are fixed and predictable.
Models Instant access to the latest proprietary models. Open source models; you choose the model and the version.
Operations The provider handles setup and maintenance. Your team handles serving, monitoring and updates.
Connectivity Requires internet access. Also runs on air-gapped networks.

Who is it for?

For technical teams that will stand up and run open source models in-house. Familiarity with Python, Linux and Docker is enough.

  • Software developers

    Mid-level and senior developers who want to learn the open source LLM ecosystem and build AI applications that run in-house.

  • Platform and infrastructure engineers

    Teams that will plan GPU capacity and build and scale the model serving layer.

  • DevOps and MLOps engineers

    Operations specialists who will own deployment, monitoring and maintenance of local LLM systems.

  • Technical leaders and architects

    Leaders choosing between cloud and on-premise based on cost, performance and security.

Which model, how much memory?

In local LLM projects the first question is hardware. Pick a model size and a precision to see how much GPU memory the model weights take. In the training you do this calculation for your own workload, including context length and the number of concurrent users.

Want this hardware installed and ready? Discover RuneBox

Model size
Precision

≈ 5.4 GBweights + 20% margin

The weights fit on a single 24 GB GPU with room left for the KV cache.

Rough estimate: weights plus a 20% margin. The ticks mark the usable 80% of each card; the rest is for the KV cache, which grows with context length and the number of concurrent requests.

Program

Seven modules. Each one is hands-on in a live GPU environment; examples are adapted to your organization's use cases.

  1. The open source LLM ecosystem

    The Qwen, Mistral, Gemma, gpt-oss, DeepSeek and GLM families; choosing a model by task, size, Turkish performance and license. GGUF and SafeTensors formats, reading a model card.

  2. Hardware planning and deployment architecture

    Estimating VRAM; how AWQ, GPTQ, FP8 and GGUF quantization affect memory and quality. Single server or cluster, tensor and pipeline parallelism across GPUs, containers with Docker.

  3. Model serving: vLLM, SGLang and Ollama

    PagedAttention and continuous batching; connecting existing applications unchanged through an OpenAI-compatible API. Quick prototypes with Ollama and llama.cpp, production with vLLM and SGLang.

    Workshop Publishing a model in-house as an OpenAI-compatible endpoint.

  4. Local embeddings and vector databases

    Multilingual embedding models such as BGE-M3 and multilingual-e5; setting up Qdrant, Milvus and pgvector on-premise and choosing between them for your use case.

  5. On-premise RAG pipeline

    Document processing for PDF, DOCX, XLSX and email, OCR for scanned documents, chunking, retrieval and reranking. Managing models and packages in air-gapped environments, network segmentation and encryption.

    Workshop An end-to-end RAG system with every component running in-house.

  6. Chatbot interface and enterprise integration

    Interfaces with Open WebUI, Chainlit and Gradio; conversation history, source citations and feedback. Authentication with SSO, integration with the intranet.

  7. Performance, cost and monitoring

    Measuring TTFT and tokens/s, load testing; optimization with batching, the KV cache and semantic caching. Monitoring with Prometheus metrics, API keys and rate limits. Total cost of ownership, cloud vs on-premise.

    Output A cloud vs on-premise cost comparison for your organization.

  8. Tailored to your organization

    The weight of each module, the models used and the exercises are agreed in a scoping call before the training, based on your team's profile and your existing infrastructure.

Get a quote

Format and certificate

Certificate

Every participant receives a RuneLab Academy certificate of attendance. On request, the training is run together with Boğaziçi University Lifelong Learning Center and the certificate is issued with BÜYEM endorsement.

Boğaziçi Üniversitesi Yaşamboyu Eğitim Merkezi logosu

The system your team builds by the end

Each layer is built in its own module. By the end, your team has built the whole system once, end to end, and keeps the setup steps, configurations and code.

Corporate network

  1. Module 6InterfaceOpen WebUIChainlitGradio
  2. Module 5RAG pipelineParsingOCRRetrievalReranker
  3. Module 4Embeddings and vector DBBGE-M3multilingual-e5Qdrantpgvector
  4. Module 3ServingvLLMSGLangOllamaOpenAI-compatible API
  5. Modules 1–2Model and hardwareQwengpt-ossGemmaAWQVRAM plan

Module 7 · Performance and monitoring

RuneBox logo

The same architecture, pre-installed: RuneBox

The ready-made version of the system you build in the training: it comes installed on your premises with NVIDIA-based GPUs, local models and RAG infrastructure, and RuneLab maintains it.

Frequently Asked Questions

What is a local LLM?

A local LLM is a language model that runs on the organization's own servers instead of a cloud API. The model is loaded onto in-house GPUs; questions, documents and answers never leave the corporate network. This training teaches, hands-on, how to build and run such a system with open source models.

How long is the training?

The length is set in a scoping call before the training, based on the scope of the modules, your team's level and whether the exercises run on your own infrastructure, and it is stated in writing in the quote.

Do we need our own GPUs for the training?

No. The GPU environment for the exercises is provided during the training; how participants reach it is planned with you beforehand, around your network restrictions. If your organization has GPU infrastructure, part of the exercises can also run on it.

Is our company data used in the exercises?

No. Exercises use sample document sets we prepare; no real company data is entered into any tool.

Which models are used?

Open model families such as Qwen, Mistral, Gemma, gpt-oss, DeepSeek and GLM. The choice depends on the task, Turkish performance, hardware and license terms. Model versions are reviewed before every training.

Does it work on air-gapped networks?

Yes. Once the models, embedding models and vector database are installed, the system runs without any outside connection. Moving models and packages into an air-gapped environment, updates and access control are covered separately in the program.

Is fine-tuning included?

This program focuses on running existing open source models on-premise. Fine-tuning with LoRA and QLoRA is planned separately when needed.

How is it different from the RAG Training?

The RAG Training focuses on designing systems that answer from corporate knowledge with cited sources, and it applies to any model, in the cloud or on-premise. This training teaches running the model itself on-premise; here RAG is built as a fully on-premise pipeline. Explore the RAG Training.

How does it relate to RuneBox?

RuneBox is the architecture taught in this training, delivered installed and maintained by RuneLab. If you want your own team to build and run the system, choose the training; if you want to start with ready-made infrastructure, choose RuneBox. Teams using RuneBox also take this training to run the system and build applications on it.

Do you issue a certificate?

Every participant receives a RuneLab Academy certificate of attendance. On request, the training can be run together with Boğaziçi University Lifelong Learning Center and the certificate issued with BÜYEM endorsement.

Who are your trainers?

The training is delivered by experienced senior AI engineers and data scientists who run enterprise AI projects at RuneLab.

Is the content adapted to our organization?

Yes. In a short call before the training we learn about your team's profile, your existing infrastructure and your use cases, and prepare the models and exercises accordingly.

How is the price set?

We prepare a quote for your organization based on the number of participants, the format and how much the content is tailored. Fill in the form and we will get back to you within one business day.

Get a quote for your team

Tell us the number of participants, your team's technical profile, and your preferred format and dates; we will get back to you within one business day.

Get a quote