DataEngBytes Sydney Overview

DataEngBytes Sydney Overview

I have just returned from my first DataEngBytes conference in Sydney. Here is a quick summary on my thoughts for the event. I plan to go into more detail on great points from presentations in later posts.

About DataEngBytes

“Run by data engineers, for data engineers” is exactly what this event was. With four in-person 2026 conferences in Brisbane, Auckland, Melbourne and Sydney, plus meetups in additional cities including Hobart and Perth, DataEngBytes definitely lives up to its About Us statement.

DataEngBytes plays an important role highlighting the work of data engineers in Australia and New Zealand and positioning the region as a global epicentre of data engineering innovation.

DataEngBytes team

As a meetup organiser myself for a number of years, my deep gratitude goes to Peter Hanssens and his team for their dedication in building and supporting a vibrant tech community.

People and Places

This is the first conference I’ve attended in Australia for a decade. The last time I was a speaker at an event on the Gold Coast. Sydney is about as far away as you can get from my three most recent conferences in San Francisco (May), Brussels (Feb) and São Paulo (Oct). I was still able to run into a fellow technologist I know from different previous events across the globe. Dr Frank Munz from Germany, a fellow Oracle ACE Director Alumnus was a presenter. I also met speakers from Brisbane, Adelaide and Perth, and attendees from many parts of Australia and even as far away as South Korea. Shout out to JiHyeok . Another fellow attendee was an ex-New Yorker like myself, and technologists I spoke with hailed from diverse locations including Russia, Turkey, Brazil and Saudi Arabia. This event drew an international audience, just like my prior American and European conferences.

Day 1 Talks - Data Engineering

Each talk brought a different perspective on data engineering with very little overlap in the scope of presentations. Speakers with different backgrounds spanning a wide variety of industries, technology products, skills and industry experience. (Except all the town-hall bankers, you know who you are.) Day 1 was dedicated to data engineers and Day 2 dedicated to AI engineers.

The 5 Vs of Big Data

Simon O’Toole from Macquarie University opened by providing us with the classic big data problem in Astronomy . Massive data ingestion requirements, immutable raw data requirements and idempotent stages of a data pipeline necessary to support global research and discovery. He left us with the question “Which of your pipelines would survive scrutiny of others?”

Jordan Simonovski gave a compelling case for wide events and a rethink of traditional observability, using segmented event types of information and the correspondingly segmented technology stack. I agree, “We need to be more exploratory from a single product” over our current OTel ways. “A metric is a query, not a place to put things” and “Every metric is a SELECT you haven’t run yet” were slides that rang true. Adding ClickStack to my reveiw new tech list.

The event moved into multi-track territory following the first break. I attended sessions with Carine Oliveira from Easygo, who spoke on the importance of Team Structure and why ownership, roles, and platform layers matter more than tools. Shilin Wu in one of two presentations from VeloDB , gave a hands-on demo of a unified real-time and AI-ready data platform, which also speaks the MySQL protocol!

After lunch, Frank Munz gave us a lesson on Spark Declarative Pipelines (SDP) putting together streaming tables an materialised views with declarative pipelines. His demo included a real-time visualisation via DataBricks of the airspace over Europe GitHub .

Bitol graduates in the Linux Foundation

The organisers were able to add to the program a great presentation by Jean-Georges “jgp” Perrin . He presented his work on the Open Data Contract Standard (ODCS) , the Linux Foundation’s Bitol project (GitHub ). This presentation showed the ODCS architecture and content, and the importance of having defined Data Contracts between contributors and consumers. In a study using Data Contracts and Data Products as the foundation for AI projects in production, JGP showed compelling results with a 3.4x improvement in maturity outcomes. Data Contracts and Data Products can underpin the lineage, context and governance, for your Data Mesh.

Congratulations to Bitol and its promotion to a graduated Linux Foundation project!

Retrieval Engineering in Luminary by Anup Sethuram

The final scheduled speaker session on day 1 was Anup Sethuram who presented Luminary (GitHub ) in his talk on Retrieval Engineering . The approach of why, and then how to fuse lexical, vector and graph searches into a single ranked result providing a better AI response. This moves beyond a traditional nearest neighbour algorithm used in vector search databases. His funnel architecture described three levels of refinement. Level 1 is candidate generation, level 2 is re-ranking, and level 3 is business logic. This all runs locally with configurable models, ensuring your data never leaves your environment. Another product for my eval list.

Day 2 Talks - AI Engineering

Kanishka Mohaia from Sonder kicked off Day 2 speaking to us on the trust gap we find with AI, including in governance and guardrails. He highlighted the need for human review in AI, drawing on his work in the health industry. He also talked about the importance of data residency and data sovereignty, and how they are not the same. He challenged us to insist that vendors be transparent about the sovereignty of your company data.

Snowflake CoCo

Travis Murphy from Snowflake introduced CoCo , a data-native AI agent harness integrated with the Snowflake platform. It includes plenty of agent skills to help you build, stream, migrate and manage data within the Snowflake ecosystem. Be sure to look at the Snowflake World Tour , offering 23 upcoming events between August and October 2026.

Vaishnavi Mohan provided us with a much deeper look and understanding of how agent memory works in Where Were We? Agentic Memory for Building Agents That Remembe . Understanding how short-term memory, long-term memory and memory scope work is a key architectural requirement for improving your agents and their interactions with our conversations.

Co-Founder of VeloDB, Matt Yi, presented his work on Apache Doris covering how this analytics and search database is designed for AI agents with the capability to provide unified data access to all enterprise data in a lakehouse implementation. Apache Doris has adopted a wide range of methods to optimise performance. Matt provided benchmarks showing the improvements over popular products including Redshift and Trino, particularly with wide tables and multi-table join queries.

OpenCode

After lunch we returned to multiple tracks, and after some room shuffling I first attended a presentation by Luke Parker sharing his work on opencode . He provided several compelling title including “If you don’t know what you want, AI won’t either!” and “Ground thagent in your reality”. As he pointed out, AI is changing every day, with many different subscriptions available. Two important simple starting points are to expand your capabilities to automate in AI, and let AI be the skeptic you need in software development, his /skeptical skill. His statement “A one shot question to AI gives you a result that is the distilled average of the Internet” reinforces that more data does not automatically provide a better answer. Some references from his talk included I Have ADHD and Theo Browne .

Rounding Errors. When 1.0 + 2.0 does not equal 3.0

Rimma Shafikova gave a great presentation on Why Exactly LLMs are Non-Deterministic and Can We Tame This Probabilistic Beast? by explaining the basic maths and the fundamental way the technology calculates results for AI, and where determinism is hard. One reason comes down to how floats are stored efficiently

Dr Jonathan Carroll entertained us with his personal exploration using a graph strategy to derive and investigate new opportunities from his personal plain markdown notes. When leveraging WikiData , some syntactic sugar in his notes, and Obsidian , possible future topics of common interest were identified from the existing relationships. Indeed a[5] does equal 5[a] (you had to be there).

Alkera CEO Rick Geo speaking at DataEngBytes in Sydney

We wrapped with another last minute guest speaker addition to the program, Rick Gao, CEO of Alkera . It is are working on ensuring AI is more aligned with the complexities of data engineering tasks, providing rich scaffolding as an operating system around a model engine.

In Conclusion

On both days Peter closed with a town hall. An opportunity for the audience to participate, prompting discussions with survey results by data practioners. This gave the opportunity for input from a variety of people, industries and roles including attendees that were not data engineers.

The final town hall slide re-iterated the key reason why people were here and what was discussed.

Everyone’s building with AI. Almost nobody says their data is ready for it. The gap isn’t the model - its the foundation beneath it.

More to come in future posts.

Tagged with: Conference Data Engineering

Related Posts

Percona Live Conference Recommendations

While many attendees are repeat offenders, if 2013 is your first MySQL conference and you are relatively new with MySQL (say < 2 years experience), it can be daunting to determine which of the 8 or more concurrent sessions you should attend during the conference.

Read more

MySQL now has two user conferences (*)

PC World has written a post with this title(*) about the upcoming MySQL Connect conference and references the Percona Live conference and an official Percona comment. As this is not syndicated in Planet MySQL I encourage you to read the full article .

Read more

MySQL conference schedule

I am one of the crazy individuals(*) that will be speaking at both the regular O’Reilly MySQL Conference and the IOUG Collaborate conference both being held in the second week of April.

Read more