10.0.0: Java 25, Python 3.13, and a new SQL engine
The platform moved to Java 25 and Tomcat 11. On top of it: Modern DataSet SQL, embedded Python 3.13 replacing Jython, streaming reads for JSON, XML and HTTP.
Etlworks 10.0.0 is the biggest release since the first one. The whole platform moved to Java 25, the application runs on Tomcat 11, and the Integration Agent, the worker you install on your own hardware, ships as a new generation with its own command line. On that foundation 10.0.0 adds a new SQL engine for extracted data, Python 3.13 in place of Jython, and streaming reads for large JSON, XML and HTTP responses.
The REST API is unchanged. Every endpoint keeps its path, parameters and responses, and existing API clients keep working.
Three changes can require action, so start there.
Your Python scripts need a review
Embedded Python moved from Jython (Python 2.7) to Python 3.13, and there is no compatibility switch. Scripts written in syntax common to both versions may run as they are. Scripts using Python 2 syntax will not: print statements, except Exception, e, xrange, iteritems, unicode, basestring.
Java classes load differently too. The Jython-style from com.example import SomeClass is gone, replaced by java.type():
import java
Jsoup = java.type("org.jsoup.Jsoup")
ArrayList = java.type("java.util.ArrayList")
Review every Python script before you upgrade the runtime that executes it. For flows that run on an Integration Agent, the executing Agent decides: a 9.x Agent still runs Jython, so an application upgrade alone does not convert anything. With mixed versions, test on the machine that actually runs the flow. The migration checklist lists the conversions in a table, and our scripting basics page has been rewritten for 3.13.
What you get for the trouble: f-strings, the bundled standard library (json, re, datetime, decimal, pathlib and the rest), importable pure-Python modules of your own, and a result that can be the final expression instead of a callback. There is still no pip in embedded Python. Package-heavy work still belongs in Execute Script Local or Remote via SSH, which runs your own interpreter and is untouched by this release.
Local SQL runs on a new engine
SQL executed locally over extracted datasets, meaning files, API responses and spreadsheets rather than a database, now runs on Modern DataSet SQL. Native database and MongoDB queries are not affected.
Most existing queries work unchanged. A query that depends on Legacy-specific syntax or behavior can fail or return a different result. Saved SQL is not rewritten, and Legacy is still there: add a directive to one statement, or select it per tenant or globally.
/* etlworks:sql-engine=Legacy */
select * from customers
The new engine is worth moving to. It adds joins between extracted relations, derived tables, independent and correlated subqueries, nonrecursive CTEs, window functions, native scalar functions and CAST, and DISTINCT aggregates. It also adds a sources namespace, which lets a query join the retained results of earlier transformations in the same flow. A CSV, a REST response and a spreadsheet can be joined in one statement, on the extracted data, before anything is loaded.
The SQL migration checklist covers the differences worth checking first, and Use SQL to extract data from non-relational sources documents the engine.
Snowflake and Databricks need one JVM option
On self-managed Java 25 deployments, the Arrow-based JDBC drivers used by Snowflake and Databricks fail until the JVM allows access to java.nio. The symptom is an InaccessibleObjectException followed by NoClassDefFoundError, in Explorer, our built-in data browser, and in flows.
--add-opens=java.base/java.nio=ALL-UNNAMED
New installers, the Agent CLI and the Docker images add it for you. Manually managed Tomcat deployments and existing Java 25 Agents running on older launchers need it added by hand, followed by a JVM restart. Hosted instances are updated by us.
Parallelism settings above 10 now take effect
Earlier versions silently limited parallel execution to 10 concurrent threads no matter what the flow was configured with. Starting with 10.0.0 the configured value is used, up to 100.
A flow saved two years ago with 64 threads ran with 10 and now runs with 64. Review the flows with high thread counts before you upgrade and confirm the sources and destinations can take the load. What changed in parallel loops. This is the change most likely to surprise someone, because nothing in the flow changed.
Streaming, in three places
Large JSON and XML. Both formats gain a Full streaming read mode. Records are read one at a time, scalar fields become columns, and nested objects and arrays are preserved as native-format string columns, ready to load into a VARIANT, JSONB or SUPER column. Memory depends on the largest single record rather than the size of the file. For JSON, combining it with the Start Node streams one specific array inside a larger document. Mapping and previews sample the source instead of reading all of it. Streaming large nested JSON and XML.
HTTP responses. Stream JSON Response reads a response incrementally, and with automatic pagination it requests the next page only once the current one has been consumed, so a large multi-page API result no longer has to fit in memory. Stream CSV Response converts JSON responses to CSV record by record for connections configured with Output as CSV.
Force Streaming, chosen for you. Eligible transformations that load a materialized source into a relational database now get Force Streaming automatically. The transformation parameter became a three-value setting: Automatic, which follows the new global and tenant policy and is the default, plus Enabled and Disabled for the cases where you want to decide. Existing flows map onto it with no rewriting, and nested sources loading into flat relational destinations are now eligible. Streaming vs Force Streaming.
A new Integration Agent, and a host CLI for the application
Agent 10.0.0 ships with a cross-platform command line for install, service management, updates, backup and restore, health checks and diagnostics. Updates are transactional: the Agent backs itself up before replacing anything and preserves your data, licenses and custom drivers. It carries a managed private Java runtime, so the host does not need one. Docker gets a plain lifecycle with data-preserving image updates.
Agent 10.0.0 requires application 10.0.0 or later. Existing 9.x Agents, back to 9.1.3, migrate through a bridge update and then one command on the host. Agents older than 10.0 keep working against the new application with their existing behavior, including the Legacy SQL engine, so the fleet does not have to move in one weekend.
The application has its own host CLI now. It installs and upgrades the full stack, downloads the licensed version directly, migrates a Java 8 installation to Java 25 with a single plan approval on Windows, upgrades PostgreSQL, takes standalone backups that include both WARs, and provides recovery and rollback commands when a migration fails. Fresh Windows installs use a pinned stack: Java 25, Tomcat 11, PostgreSQL 17. Upgrades preserve settings, data and drivers.
Tomcat 9 deployments are not supported on 10.0.0. The Java 8 application can run alongside the new engine during a staged migration.
Revert and restore from the CLI
The application CLI gained eleven commands for getting things back.
revert-flow, revert-connection, revert-format, revert-schedule, revert-agent and revert-macro restore any saved revision of an artifact that still exists. Pair them with the matching *-history command, which gives you the revision UUID to pass:
flow-history 123
revert-flow 123 973d824c-3d0f-4d9f-8a56-decf8f08ebce
list-deleted-flows and its four siblings show what is in the Recycle Bin, and restore-flow and its four siblings bring a deleted artifact back by name or ID, recovering supported missing dependencies along with it.
Both families save immediately, require a tenant Administrator or Super Admin, and use the standard CLI confirmation challenge, with --force for unattended automation. The CLI reference documents the rules that bite, including the fact that revert takes an exact revision UUID and that current and last remain diff-only selectors.
Smaller things worth knowing
Two administrative gaps closed. A tenant administrator or super admin can now disable two-factor authentication for a locked-out user directly from the user editor; the action is permission-checked on the server, audited, and confirmed by email to the user, who re-enrolls. And the Update existing spreadsheet option for XLSX now works when the destination is remote storage: FTP and FTPS, SFTP, SMB Share, WebDAV, Amazon S3, Azure Storage, Google Cloud Storage, Google Drive, SharePoint, OneDrive, Dropbox and Box. Other worksheets, formulas and styles survive the update.
The engine itself allocates less and releases sooner. Rows are built at their known width, repeated SQL statements and loop scripts are prepared once per execution, dataset rows are released as soon as the last consumer finishes, and parallel loops keep bounded iteration queues so memory stops growing with iteration count. Stop requests are picked up quickly. A new per-execution worker budget, etl.concurrency.max.workers, defaults to 100 and prevents deadlocks in deeply nested parallel flows. Sorting on numeric and date keys, date format detection and common string conversions are all faster.
JSON reads with a Start Node now select the node in a single pass rather than loading the whole document, controlled by Optimize Start Node selection and on by default. XML reading, writing, XSLT and XQuery use less memory, and repeated XSLT transformations reuse the compiled stylesheet. Excel processing is faster and always releases workbooks and temporary files.
Flow Findings, the view where our static analysis flags known good and bad practices in a flow, now bases findings on what the configured connection, format and mapping actually do instead of on format labels, names the transformation, field or column it refers to, and links a finding inside a nested flow back to the flow that owns it. The execution plan reports the effective bind-variable and batching settings next to the configured ones and explains any adjustment.
Version labels across the application, the Agent and the APIs now carry the runtime, as in 10.0.0 (Java 25), which makes a mixed fleet auditable during a migration.
10.0.0 also carries everything from 9.9.10 through 9.9.12: Kerberos authentication for SMB Share and Azure Files, FTP and FTPS transfer mode with a pre-transfer SITE command that completes the z/OS mainframe support, MongoDB CDC authentication and filtering options, Excel worksheets as a tree in Explorer, and version history reachable from the Connections, Formats and Listeners lists. The fixes from those builds are included too, among them MongoDB CDC capture scope and capture target being applied rather than silently ignored.
Upgrading
Read the full release notes first, then the on-premise migration guide for the application and the Agent migration guide for the fleet. Hosted customers are migrated by us.
If you run embedded Python, start the script review now rather than at upgrade time. It is the one part of this release that needs a person to look at code, and the migration checklist makes it mechanical.