Models

Open-weight and open-source models.

Curated and configured for production inference. Exact model availability, tokenizer behavior, context limits, supported parameters, and deployment configuration are provided during onboarding.

Deployable model families

Supported model families may include Qwen, DeepSeek-derived distilled models, Mistral, Gemma, and other open-weight or commercially deployable models, subject to license review, hardware fit, and deployment-specific approval. Exact model IDs are provided during onboarding.

Qwen

Alibaba

General-purpose LLM

DeepSeek

DeepSeek

Reasoning · efficient

Mistral

Mistral AI

Open-weight · performant

Gemma

Google

Lightweight · efficient

Llama

Meta

Open-weight · widely supported

Kimi

Moonshot

Long-context

Additional open-source model families available on request.

Services

Inference services for production AI teams.

Three deployment models share the same operational discipline — select the deployment shape that matches your security, performance, region, and cost requirements.

Hosted Open-Weight Model Inference

Access production-ready open-weight and open-source models through an OpenAI-compatible API. Integrate with familiar SDKs while reducing provider lock-in and improving cost control.

  • OpenAI-compatible API endpoints
  • Access to open-weight and open-source LLMs
  • Model benchmarking and selection
  • Token cost optimization
  • Usage visibility and performance monitoring

Private and On-Prem Model Operations

Deploy and operate models inside customer-controlled environments for teams with strict security, compliance, privacy, or data residency requirements.

  • On-prem or private cloud deployment
  • Model serving setup and management
  • Monitoring and operational support
  • Cost and latency tuning
  • Data control and deployment isolation

Dedicated Managed Inference

Run dedicated model deployments managed by Latens for predictable performance, workload isolation, and production-grade reliability.

  • Dedicated capacity
  • Fully managed model serving
  • Deployment-specific regional options
  • Performance and latency optimization
  • Custom operational support
Tokenization

Token accounting

Latens counts token usage with the model-native tokenizer used by each deployed model. All deployed models use llama.cpp tokenization. Prompt tokens are counted after applying the model's chat template and special-token rules. Completion tokens are counted from generated token IDs.

Data handling

Zero Data Retention production mode

For ZDR production endpoints, Latens does not intentionally store prompt text or completion text after request completion and does not use customer inputs or outputs for training, fine-tuning, evaluation, benchmarking, or product improvement without separate written authorization.

Operations

Production SLO and support

Covered dedicated production inference deployments are operated against a 99.95% monthly availability SLO. Covered production incidents receive a 6-hour initial response Support SLA through the designated support contact. The availability SLO is an operational target unless a signed agreement states otherwise.

Notice

Data & operations notice

Latens provides dedicated AI inference infrastructure for business customers and approved platform integrations. The information below summarizes how Latens handles production inference data, token accounting, deployment configuration, and operational support.

Production inference data

Production ZDR endpoints do not intentionally store prompt text, completion text, or persistent result caches after request completion. Customer inputs and outputs are not used for model training, fine-tuning, evaluation, benchmarking, product improvement, or marketing unless separately authorized in writing.

Token accounting

Token usage is counted using the model-native tokenizer for each deployed model. Latens uses llama.cpp tokenization for deployed models. Prompt tokens are counted after chat-template rendering and special-token handling. Completion tokens are counted from generated token IDs.

Deployment model

Latens supports dedicated inference deployments across agreed deployment regions. Region, model availability, context limits, supported parameters, tokenizer behavior, data-handling mode, and deployment-specific configuration are confirmed during onboarding.

Operational commitments

Covered dedicated production inference deployments are operated against a 99.95% monthly availability SLO. Covered production incidents receive a 6-hour initial response Support SLA through the designated support contact. The availability SLO is an operational target unless a signed agreement states otherwise.

Contact

General inquiries: contact@latens-ai.com