yoDEV Decisions: compara Jev con GPT, Gemini, Claude o cualquier LLM con tus propios datos

yoDEV Decisions: Compare Jev with GPT, Gemini, Claude, or any LLM using your own data

yoDEV Decisions: compara Jev con GPT, Gemini, Claude o cualquier LLM con tus propios datos

Today we are launching yoDEV Decisions, a free tool for yoDEV members that runs the same prompt on Jev and the LLM of your choice, row by row, and compares the results: accuracy, calibration, cost, and latency. The idea came from our article on TypeSafe’s Jev. We concluded there that a provider’s numbers are no substitute for your own measurements, and that teams in Ibero-America need to measure with their own data, in their own language. Decisions is that measurement. yoDEV covers the calls to Jev. For the LLM side, you connect your own OpenRouter account.

Hoy abrimos yoDEV Decisions, una herramienta gratuita para miembros de yoDEV que ejecuta la misma pregunta en Jev y en el LLM que elijas, fila por fila, y compara los resultados: exactitud, calibración, costo y latencia. La idea surgió de nuestro artículo sobre Jev de TypeSafe. Allí concluimos que las cifras de un proveedor no sustituyen una medición propia y que los equipos de Iberoamérica necesitan medir con sus datos, en su idioma. Decisions es esa medición. yoDEV cubre las llamadas a Jev. Para el lado del LLM conectas tu propia cuenta de OpenRouter.

What do the first results show? For the public results, we chose four popular LLMs (GPT-4o mini, Gemini 2.5 Flash, Claude Haiku 4.5, and Llama 3.3 70B) and four public datasets, with 200 rows each, with the prompt formulated in Spanish. With your own data, you can compare Jev with any model available on OpenRouter, including free models.

¿Qué muestran los primeros resultados? Para los resultados públicos elegimos cuatro LLM populares (GPT-4o mini, Gemini 2.5 Flash, Claude Haiku 4.5 y Llama 3.3 70B) y cuatro conjuntos de datos públicos, de 200 filas cada uno, con la pregunta formulada en español. Con tus propios datos puedes comparar Jev con cualquier modelo disponible en OpenRouter, incluidos modelos gratuitos.

DatasetTaskJevBest LLM
Banking77Bank customer intent (77 options)82.0%84.0% (Gemini 2.5 Flash)
SMS spamIs it spam?97.0%98.0% (Gemini 2.5 Flash)
Spanish moderationIs it offensive?98.0%98.5% (GPT-4o mini & Llama 3.3 70B)
Spanish sentimentPositive, neutral, or negative66.0%75.0% (Gemini 2.5 Flash)
ConjuntoTareaJevMejor LLM
Banking77Intención de un cliente bancario (77 opciones)82,0 %84,0 % (Gemini 2.5 Flash)
SMS spam¿Es spam?97,0 %98,0 % (Gemini 2.5 Flash)
Moderación en español¿Es ofensivo?98,0 %98,5 % (GPT-4o mini y Llama 3.3 70B)
Sentimiento en españolPositivo, neutral o negativo66,0 %75,0 % (Gemini 2.5 Flash)

Three observations: Latency: Jev responded with a median of about 250ms across all four sets. The four LLMs took between 1 and 2.3 seconds. Cost: Jev was the cheapest in all four sets. In Banking77, it cost about 0.08 USD per 1,000 rows, compared to 0.28 USD for Gemini 2.5 Flash and 1.84 USD for Claude Haiku 4.5. Accuracy: Jev doesn’t win everything. In intent classification, spam, and moderation, it stays within two points or less of the best LLM, and even ahead of some. In Spanish sentiment, it stays nine points behind. That last result is the reason for the tool’s existence: an average doesn’t tell you how a model behaves on your specific task.

Tres observaciones: Latencia: Jev respondió con una mediana de unos 250 ms en los cuatro conjuntos. Los cuatro LLM tardaron entre 1 y 2,3 segundos. Costo: Jev fue el más barato en los cuatro conjuntos. En Banking77 costó unos 0,08 USD por cada 1.000 filas, frente a 0,28 USD de Gemini 2.5 Flash y 1,84 USD de Claude Haiku 4.5. Exactitud: Jev no gana en todo. En clasificación de intención, spam y moderación queda a dos puntos o menos del mejor LLM, e incluso por delante de algunos. En sentimiento en español queda nueve puntos por detrás. Ese último resultado es la razón de ser de la herramienta: un promedio no te dice cómo se comporta un modelo en tu tarea.

What happens if you combine Jev with an LLM? Decisions also calculates a combined strategy: Jev responds when its confidence exceeds a threshold, and the remaining rows go to the LLM. The tool recommends that threshold based on the data. In Spanish sentiment, that combination reached 74.5% accuracy while sending only 40.5% of the rows to Gemini. You get almost the accuracy of the LLM, with Jev resolving the majority of the rows at its speed and cost.

¿Qué pasa si combinas Jev con un LLM? Decisions también calcula una estrategia combinada: Jev responde cuando su confianza supera un umbral y el resto de filas pasa al LLM. La herramienta recomienda ese umbral a partir de los datos. En sentimiento en español, esa combinación alcanzó un 74,5 % de exactitud enviando a Gemini solo el 40,5 % de las filas. Obtienes casi la exactitud del LLM, con Jev resolviendo la mayoría de las filas a su velocidad y su costo.

How do I test it with my data? Log in to decisions.yodev.dev with your yoDEV account. If you don’t have one yet, registration is free. Connect OpenRouter for the LLM side and choose any of its models. You pay for those calls with your own balance, or choose a free model. Your OpenRouter key is saved in your browser. Choose a predefined set or upload a CSV with a text column and a label column with the expected answer: up to 2,000 rows. Review the results: accuracy, calibration, cost, latency, the rows where the models don’t match, and the recommended combined strategy. Each run is automatically saved to your yoDEV Workplace. You can phrase the prompt in English and Spanish to measure the difference. Jev was trained primarily in English.

¿Cómo lo pruebo con mis datos? Inicia sesión en decisions.yodev.dev con tu cuenta de yoDEV. Si aún no tienes una, el registro es gratuito. Conecta OpenRouter para el lado del LLM y elige cualquiera de sus modelos. Pagas esas llamadas con tu propio saldo, o eliges un modelo gratuito. Tu clave de OpenRouter se guarda en tu navegador. Elige un conjunto predefinido o sube un CSV con una columna text y una columna label con la respuesta esperada: hasta 2.000 filas. Revisa los resultados: exactitud, calibración, costo, latencia, las filas en las que los modelos no coinciden y la estrategia combinada recomendada. Cada ejecución se guarda automáticamente en tu yoDEV Workplace. Puedes formular la pregunta en inglés y en español para medir la diferencia. Jev se entrenó principalmente en inglés.

What about the data I upload? Only the metrics and row numbers are saved to the Workplace, never the text of your rows. The rows you upload are deleted 30 days after your last run with that set, and you can delete them earlier at any time.

¿Qué pasa con los datos que subo? Al Workplace solo se guardan las métricas y los números de fila, nunca el texto de tus filas. Las filas que subes se eliminan 30 días después de tu última ejecución con ese conjunto, y puedes eliminarlas antes en cualquier momento.

What about OpenAI’s Decisions API? On September 29, 2026, OpenAI announced a Decisions API at their DevDay: a model that chooses between predefined responses by the developer, instead of generating text—the same category as Jev. As of September 30, it is in limited preview, with no public API reference, pricing, or limits.

¿Qué hay de la Decisions API de OpenAI? El 29 de septiembre de 2026, OpenAI anunció en su DevDay una Decisions API: un modelo que elige entre respuestas predefinidas por el desarrollador, en lugar de generar texto, la misma categoría que Jev. Al 30 de septiembre está en vista previa limitada, sin referencia pública de la API, precios ni límites.

Summary of DevDay 2026: That two providers are betting on the same idea confirms that constrained decisions are becoming a distinct piece of AI applications. It also makes the age-old question more important: which one works best for your task? When that API has public access, it makes sense to measure it with the same method.

Resumen del DevDay 2026: Que dos proveedores apuesten por la misma idea confirma que las decisiones acotadas se están convirtiendo en una pieza propia de las aplicaciones con IA. También vuelve más importante la pregunta de siempre: ¿cuál funciona mejor en tu tarea? Cuando esa API tenga acceso público, tiene sentido medirla con el mismo método.

What are the limits of these measurements? They are 200 rows per set and a single run per model, performed on September 29 and 30, 2026, with Jev 1.13. Latency depends on where it is measured from. Measure from your own region before drawing conclusions. The accuracy of Banking77 and SMS spam is measured with English texts; moderation and sentiment with Spanish texts. The moderation set contains offensive language. Each results page links to the source and license of its dataset. yoDEV Decisions is an independent tool from yoDEV, with no affiliation with TypeSafe AI or OpenAI. Jev is a product of TypeSafe AI.

¿Qué límites tienen estas mediciones? Son 200 filas por conjunto y una sola ejecución por modelo, realizadas el 29 y 30 de septiembre de 2026 con Jev 1.13. La latencia depende de desde dónde se mide. Mide desde tu propia región antes de sacar conclusiones. La exactitud de Banking77 y SMS spam se mide con textos en inglés; la de moderación y sentimiento, con textos en español. El conjunto de moderación contiene lenguaje ofensivo. Cada página de resultados enlaza a la fuente y a la licencia de su conjunto de datos. yoDEV Decisions es una herramienta independiente de yoDEV, sin afiliación con TypeSafe AI ni con OpenAI. Jev es un producto de TypeSafe AI.