Software development, artificial intelligence and automation.
This site gathers my work, hands-on research and experience in software development, applied artificial intelligence and technical infrastructure.
Knowledge areas
I work at the intersection of development, infrastructure and artificial intelligence.
Projects
Personal work and research in software development and artificial intelligence.
Booking app for restaurants with a conversational AI assistant. Multi-tenant with restaurant and customer roles, integrated management panel, search by cuisine and area, and personalised recommendations. Multilingual support (ES/EN).
Experimental merge of Qwen3.8-Flash-Next (60% Swift1.5 + 40% Tinfield-1), optimized for coding, tool-calling and efficient reasoning. Quantized in GGUF Q4/Q8, with vision support and local execution via llama.cpp.
Rust daemon for Linux that dynamically regulates the power limit of NVIDIA GPUs based on their temperature. Designed for AI inference servers with sustained loads and multi-GPU setups.
Membership management system for the Tibi municipal swimming pool. Public registration with card payment (Redsys), QR scanner for entries and exits, admin panel with roles, history with CSV export and editable legal pages. Custom MVC in PHP with no external framework.
ERP Academias
Custom multi-tenant ERP for exam-prep academies. Records management, KPI dashboard, historical snapshot system with data integrity, public API with token authentication, Stripe integration, and migrations from other systems.
On-premise AI system
Private enterprise ChatGPT with RAG, local models and no data leakage to external providers. Private infrastructure with 2× RTX A5000.
Question Bank
Quiz management platform for exam-prep academies. Multiple-choice questions, categorization by topic and tags, difficulty levels. JWT authentication, full REST API with Swagger UI. Dual DB: MongoDB (Beanie ODM) and MariaDB (SQLAlchemy) interchangeable via factory pattern.
LLaMA Server GUI
GTK4 graphical interface for managing LLaMA servers locally. GGUF model control, inference parameter configuration and system resource monitoring (CPU/RAM/GPU) with psutil.
Internal projects
Business process automation, API and legacy system integration. Workflows with AI agents and MCP tools.
CLI Agent
Conversational AI agent in the terminal with a ReAct architecture. Multi-provider (Ollama, Claude, Gemini, OpenAI, DeepSeek). 6 built-in tools: command execution, filesystem, code editor, web searches, deep thinking and sequential thinking. Automatic context management with smart compaction.
Technical articles
Notes on development, AI and infrastructure.
How I built TakeOnMe 3.8 Flash Next: a local merge for programming, agents and tools
I needed a single local model capable of reasoning, coding, working in the terminal and using tools without switching between several. TakeOnMe 3.8 Flash Next is a 60/40 merge of Swift 1.5 and Tinfield over Qwen3.8-Flash-Next, quantized in mixed Q4/Q8 GGUF.
NVIDIA CMP 170HX: from mining cards to 104 GB of VRAM for local AI
I bought two NVIDIA CMP 170HX mining cards (8 GB and 10 GB) and the unlock left me with 64 + 40 = 104 GB of GA100 VRAM. How the unlock works, how they behave inferencing LLMs with llama.cpp and when they make sense.
Qwen3.8 27B vs Ornith 1.5 35B A3B
I have spent hours testing Qwen3.8 27B and Ornith 1.5 35B-A3B working with code and agents. Two very capable models, but with surprisingly different ways of working. A comparison based on real usage experience, beyond what the benchmarks tell.
vLLM vs llama.cpp: which server to use for local inference?
A comparison between vLLM and llama.cpp for serving language models locally: GPU performance, GGUF and FP8 quantization, multi-GPU and when to choose each engine.
The impact of long context and agents on the cost of AI
Artificial intelligence is increasingly powerful, but also more expensive to use. The rise in bills in 2026 is not only due to price increases by providers, but to deep changes in how modern models are used.
Artificial intelligence applied to real software development
Artificial intelligence is no longer a future promise: it is a practical tool for building, maintaining and improving business applications.
A practical guide to GGUF quantization for local inference
How to optimize language models through GGUF quantization to run them on local hardware with llama.cpp and vLLM.
Setting up an inference server with Ollama and vLLM locally
Steps to deploy a local inference server with Ollama and vLLM using NVIDIA GPUs and open models.
AI agents with MCP: practical automation for businesses
An exploration of AI agents using the MCP protocol and how to apply them to real business automation.