java coding – Libera #java

Symptoms and Solutions

I casually watch a lot of forums related to Java: IRC, of course, and Discord, and Slack, and Reddit. On Reddit, on r/java, there’re some occasionally quite interesting projects offered for consideration; they’re usually pet projects for the authors, and that’s healthy and useful for everyone.

However, there’s also a low hum of … tooling. Two really interesting projects showed up recently, for example, that avoided standard tooling: one used PowerShell, another used a bash script to build.

Both of these projects work. They both generated the expected output; the bash script actually went way above and beyond (replicating what a good Gradle script might do), downloading a JVM and manually grabbing the dependencies via wget, then manually calling javac and jlink in a manner that would make Mark Reinhold proud, then building an AppImage for x86_64 using the platform-appropriate tools.

The Reddit discussion on that latter one did not go especially well, however, because the tooling discussion ended up front and center, with what looks like a lot of hurt feelings.

“This looks like a novice project” was the import of one comment, with another being that the tooling might have helped; that went south quickly. It was not a “novice project.” The author was solving a problem they encountered in their own way, including implementing some X11 protocols in Java. Project infrastructure choices don’t make one a “novice,” especially as an insult.

The thing is, the comment about tooling might be pretty apt. My personal thought about the project, based on the discussion, was “why not tooling?” Are the tools insufficient? The project owner built an x86_64 image specifically; that may suit their purposes (and as the primary project consumer, their purposes are most important) but what would it have taken to build an ARM image? Would the tooling have been sufficient for that?

It might not have been. The way I see it, there’re a few reasons the tooling would be avoided:

Lack of awareness. The author said they knew about the tooling and avoided it, but maybe they didn’t check to see if the tooling could create AppImages.
Lack of functionality. Maybe they did check, but the functionality wasn’t there, or wasn’t sufficient.
Lack of documentation. Maybe the functionality was there, but was documented poorly enough that it was easier to do it the straightforward way instead. (Is that OSGi I hear weeping in the distance? It just might be.)
Personal preference. Maybe the OP knew about the tooling and just didn’t care.

If personal preference is involved, well, the discussion’s over. At that point, it’s like arguing that the author’s favorite number should clearly be 18 instead of 11. It’s their project, not anyone else’s, and if they make a decision based on their preference, that’s their right and power.

But the other options, well, those are things that we as a community might be able to help. And I think we should consider helping, even if the “help” might not apply to this specific circumstance.

So: We can help, but we need to consider:

How we approach the help
What use cases are poorly served by the tooling today
How are we evangelizing our practices, and why
What would “better” look like?
Do we even understand each other to understand why people choose different approaches?

These aren’t necessarily idle musings. We want Java to remain vital, we want to do what we can to establish common practices so we can help each other, we want to make sure that we’re not presenting barriers to people whose skills we want in our ecosystem; it’d be nice if we could offer commentary without sneering (and, to be fair, accept commentary without bristling.)

Let’s all be kind to each other, an effort that would make us more effective as well.

Locale-specific numbers

One of our channel members mentioned a failure in parsing “-1” to an int given a Locale. That’s … fascinating, actually, as I (dreamreal) was unaware of any numbering systems under which that would fail. So, as any programmer would, I wanted a program to show me the locales for which “-1” wasn’t the proper representation for “negative one.”

Being a programmer, I … immediately ran to an AI (Claude, specifically) and had it generate a table of the possibilities, along with the associated locales, that did not fit the “normal” representation of “-1” – where, well, my locale (US English) was “normal.”

If you’re not English, please recognize the humor here – I’m well aware that Urdu readers would think their representation was “normal” and “-1” was not. Or, well, I am aware now and wasn’t before, and my use of “normal” is entirely meant to poke fun at my own English-centric expectations, because it didn’t occur to me that negatives using Arabic sigils might not be the same everywhere.

Anyway, this is what came out of it, for my Java 25 installation. The table shows the representation (the best that I could get WordPress to trivially display it, at least), the Unicode transcription, and then a list of the locales that emitted that particular representation. I didn’t bother including “-1” because, well, it’s a ginormous list and this table is long enough already.

Representation	Character Analysis	Locales (168 total)
`-?`	`len:2, chars:U+002D,U+07C1`	nqo (N’Ko), nqo_GN (N’Ko (Guinea)), nqo_GN_#Nkoo (N’Ko (N’Ko, Guinea))
`-?`	`len:2, chars:U+002D,U+0967`	bgc (Haryanvi), bgc_IN (Haryanvi (India)), bgc_IN_#Deva (Haryanvi (Devanagari, India)), bho (Bhojpuri), bho_IN (Bhojpuri (India)), bho_IN_#Deva (Bhojpuri (Devanagari, India)), mr (Marathi), mr_IN (Marathi (India)), mr_IN_#Deva (Marathi (Devanagari, India)), ne (Nepali), ne_IN (Nepali (India)), ne_NP (Nepali (Nepal)), ne_NP_#Deva (Nepali (Devanagari, Nepal)), raj (Rajasthani), raj_IN (Rajasthani (India)), raj_IN_#Deva (Rajasthani (Devanagari, India)), sa (Sanskrit), sa_IN (Sanskrit (India)), sa_IN_#Deva (Sanskrit (Devanagari, India))
`-?`	`len:2, chars:U+002D,U+09E7`	as (Assamese), as_IN (Assamese (India)), as_IN_#Beng (Assamese (Bangla, India)), bn (Bangla), bn_BD (Bangla (Bangladesh)), bn_BD_#Beng (Bangla (Bangla, Bangladesh)), bn_IN (Bangla (India)), mni (Manipuri), mni_IN (Manipuri (India)), mni_IN_#Beng (Manipuri (Bangla, India)), mni__#Beng (Manipuri (Bangla))
`-?`	`len:2, chars:U+002D,U+0E51`	th_TH_TH_#u-nu-thai (Thai (Thailand, TH, Thai Digits))
`-?`	`len:2, chars:U+002D,U+0F21`	dz (Dzongkha), dz_BT (Dzongkha (Bhutan)), dz_BT_#Tibt (Dzongkha (Tibetan, Bhutan))
`-?`	`len:2, chars:U+002D,U+1041`	my (Burmese), my_MM (Burmese (Myanmar (Burma))), my_MM_#Mymr (Burmese (Myanmar, Myanmar (Burma)))
`-?`	`len:2, chars:U+002D,U+1C51`	sat (Santali), sat_IN (Santali (India)), sat_IN_#Olck (Santali (Ol Chiki, India)), sat__#Olck (Santali (Ol Chiki))
`-?`	`len:3, chars:U+061C,U+002D,U+0661`	ar_BH (Arabic (Bahrain)), ar_DJ (Arabic (Djibouti)), ar_EG (Arabic (Egypt)), ar_EG_#Arab (Arabic (Arabic, Egypt)), ar_ER (Arabic (Eritrea)), ar_IL (Arabic (Israel)), ar_IQ (Arabic (Iraq)), ar_JO (Arabic (Jordan)), ar_KM (Arabic (Comoros)), ar_KW (Arabic (Kuwait)), ar_LB (Arabic (Lebanon)), ar_MR (Arabic (Mauritania)), ar_OM (Arabic (Oman)), ar_PS (Arabic (Palestinian Territories)), ar_QA (Arabic (Qatar)), ar_SA (Arabic (Saudi Arabia)), ar_SD (Arabic (Sudan)), ar_SO (Arabic (Somalia)), ar_SS (Arabic (South Sudan)), ar_SY (Arabic (Syria)), ar_TD (Arabic (Chad)), ar_YE (Arabic (Yemen)), sd (Sindhi), sd_IN (Sindhi (India)), sd_PK (Sindhi (Pakistan)), sd_PK_#Arab (Sindhi (Arabic, Pakistan)), sd__#Arab (Sindhi (Arabic))
`-1`	`len:3, chars:U+200E,U+002D,U+0031`	ar (Arabic), ar_001 (Arabic (world)), ar_AE (Arabic (United Arab Emirates)), ar_DZ (Arabic (Algeria)), ar_EH (Arabic (Western Sahara)), ar_LY (Arabic (Libya)), ar_MA (Arabic (Morocco)), ar_TN (Arabic (Tunisia)), he (Hebrew), he_IL (Hebrew (Israel)), he_IL_#Hebr (Hebrew (Hebrew, Israel)), ur (Urdu), ur_PK (Urdu (Pakistan)), ur_PK_#Arab (Urdu (Arabic, Pakistan))
`-?`	`len:4, chars:U+200E,U+002D,U+200E,U+06F1`	ks (Kashmiri), ks_IN (Kashmiri (India)), ks_IN_#Arab (Kashmiri (Arabic, India)), ks__#Arab (Kashmiri (Arabic)), lrc (Northern Luri), lrc_IQ (Northern Luri (Iraq)), lrc_IR (Northern Luri (Iran)), lrc_IR_#Arab (Northern Luri (Arabic, Iran)), mzn (Mazanderani), mzn_IR (Mazanderani (Iran)), mzn_IR_#Arab (Mazanderani (Arabic, Iran)), pa_PK_#Arab (Punjabi (Arabic, Pakistan)), pa__#Arab (Punjabi (Arabic)), ps (Pashto), ps_AF (Pashto (Afghanistan)), ps_AF_#Arab (Pashto (Arabic, Afghanistan)), ps_PK (Pashto (Pakistan)), ur_IN (Urdu (India)), uz_AF_#Arab (Uzbek (Arabic, Afghanistan)), uz__#Arab (Uzbek (Arabic))
`??`	`len:3, chars:U+200E,U+2212,U+06F1`	fa (Persian), fa_AF (Persian (Afghanistan)), fa_IR (Persian (Iran)), fa_IR_#Arab (Persian (Arabic, Iran))
`-?`	`len:3, chars:U+200F,U+002D,U+0661`	ckb (Central Kurdish), ckb_IQ (Central Kurdish (Iraq)), ckb_IQ_#Arab (Central Kurdish (Arabic, Iraq)), ckb_IR (Central Kurdish (Iran))
`?1`	`len:2, chars:U+2212,U+0031`	et (Estonian), et_EE (Estonian (Estonia)), et_EE_#Latn (Estonian (Latin, Estonia)), eu (Basque), eu_ES (Basque (Spain)), eu_ES_#Latn (Basque (Latin, Spain)), fi (Finnish), fi_FI (Finnish (Finland)), fi_FI_#Latn (Finnish (Latin, Finland)), fo (Faroese), fo_DK (Faroese (Denmark)), fo_FO (Faroese (Faroe Islands)), fo_FO_#Latn (Faroese (Latin, Faroe Islands)), gsw (Swiss German), gsw_CH (Swiss German (Switzerland)), gsw_CH_#Latn (Swiss German (Latin, Switzerland)), gsw_FR (Swiss German (France)), gsw_LI (Swiss German (Liechtenstein)), hr (Croatian), hr_BA (Croatian (Bosnia & Herzegovina)), hr_HR (Croatian (Croatia)), hr_HR_#Latn (Croatian (Latin, Croatia)), ksh (Colognian), ksh_DE (Colognian (Germany)), ksh_DE_#Latn (Colognian (Latin, Germany)), lt (Lithuanian), lt_LT (Lithuanian (Lithuania)), lt_LT_#Latn (Lithuanian (Latin, Lithuania)), nb (Norwegian Bokmål), nb_NO (Norwegian Bokmål (Norway)), nb_NO_#Latn (Norwegian Bokmål (Latin, Norway)), nb_SJ (Norwegian Bokmål (Svalbard & Jan Mayen)), nn (Norwegian Nynorsk), nn_NO (Norwegian Nynorsk (Norway)), nn_NO_#Latn (Norwegian Nynorsk (Latin, Norway)), no (Norwegian), no_NO (Norwegian (Norway)), no_NO_#Latn (Norwegian (Latin, Norway)), no_NO_NY (Norwegian (Norway, Nynorsk)), rm (Romansh), rm_CH (Romansh (Switzerland)), rm_CH_#Latn (Romansh (Latin, Switzerland)), se (Northern Sami), se_FI (Northern Sami (Finland)), se_NO (Northern Sami (Norway)), se_NO_#Latn (Northern Sami (Latin, Norway)), se_SE (Northern Sami (Sweden)), sl (Slovenian), sl_SI (Slovenian (Slovenia)), sl_SI_#Latn (Slovenian (Latin, Slovenia)), sv (Swedish), sv_AX (Swedish (Åland Islands)), sv_FI (Swedish (Finland)), sv_SE (Swedish (Sweden)), sv_SE_#Latn (Swedish (Latin, Sweden))

List.remove() oddities

From reddit’s r/java, a user observed that List.remove(Object) removes the first object whose equals() method returns true. Thus, if two objects have object equality with each other, but you wish to remove the second, you’re going to have to find it and return it with the indexed version of remove(int) instead of the remove(Object) call.

yawkat observed that you can use Collection.removeIf() (which List inherits) and get it done as well:

list.removeIf(o -> o == objectToRemove);

It’d probably still be better to fix equals() so it was more specific, but sometimes your code doesn’t always fit the circumstances you want.

WebMvcConfigurationSupport Mangles ISO8601 Timestamps in Spring Boot

I was working on a test in my Spring-Boot app and noticed that the timestamps in JSON output were sometimes formatted incorrectly. Luckily I was able to identify the issue and fix it.

During an integration test I noticed that an expected JSON output of "2019-08-26T08:22:21Z" was actually 1566807741.000000000 even though in other tests in my application the automatic conversion from java.time.Instant to an ISO8601-formatted string was working just fine. As is the default for Spring-Boot, Jackson’s ObjectMapper was used for the conversion. So what could make Jackson work correctly in one case but not another?

First I thought that I did something wrong with configuring so I did a lot of googling in order to find out what I was doing wrong. I added jackson-datatype-jsr310 to my project’s dependencies, I disabled Jackson’s “write dates as timestamps” serializer feature in four different ways but the error would persist. I dug into how the ObjectMapper used by Spring was created; maybe the feature was not disabled correctly? No, it was disabled just fine but Jackson still would not format the Instant properly.

During my research on how to create your own ObjectMapper and have Spring use it, I stumbled upon this question. It had nothing to do with my immediate problem but it mentioned that when using a @Configuration class extending WebMvcConfigurationSupport in conjunction with @EnableWebMvc Spring-Boot’s auto-configuration would be disabled. A disabled auto-configuration could explain an ObjectMapper that can not format Instants properly!

As it turns out, I had a SwaggerConfiguration class that extended WebMvcConfigurationSupport. Following the hints in the StackOverflow post I replaced the superclass by WebMvcConfigurerAdapter only to discover that it was deprecated. Fixing that was easy, though, as the interface WebMvcConfigurer could (and should!) be used instead. And this made my integration test work…

…although at this point I am still not 100% sure why. I am not using @EnableWebMvc anywhere (just @SpringBootApplication) so I do not yet know why the SwaggerConfiguration class extending WebMvcConfigurationSupport would disable Spring-Boot’s auto-configuration mechanisms. Spring’s usually excellent documentation is rather unclear as to what happens if you only do one of those things.

I also don’t know why the Jackson mapper in the working test had its “write dates as timestamps” serializer feature disabled but the one in the other test had not. From my understanding it must have been disabled so that Jackson’s jackson-datatype-jsr310 module was used—it was already a part of my dependencies without me knowing so explicitely adding it to my project’s build file did not actually change anything.

And even though my mental model of some things involving Spring has been improved by getting this to work I will not investigate this any further. My integration tests work and that shall be good enough for me, for today.

Jackson-databind and Default Typing Vulnerabilities

Today, GitHub sent out security notices to owners of projects using old jackson-databind versions (older than 2.8.11.1 and 2.9.5). These notices pertain to this issue. I have talked about its relevance before on IRC, but since it is getting more attention now, I will describe it here again.

The Problem

The “bug” comes from using the so-called default typing. This feature allows a user to deserialize subclasses (or even Object) without specifying the full possible type hierarchy. Consider this model:

interface I {
}
@Data
class A implements I {
    private int i;
}
@Data
class B implements I {
    private boolean b;
}

Now, if we wish to serialize the interface I, you usually need to specify some sort of type info. This is typically done through an annotation on the interface:

@JsonTypeInfo(use = JsonTypeInfo.Id.NAME)
@JsonSubTypes({@JsonSubTypes.Type(value = A.class, name = "A"), @JsonSubTypes.Type(value = B.class, name = "B")})
interface I {
}

If we now serialize an object of type I, we get a result like this:

new ObjectMapper().writerFor(I.class).writeValueAsString(new A())
{"@type":"A","i":0}

Deserializing this JSON works as expected.

Default typing

This annotation-based registration is fine for small use cases, but can get cumbersome if the types are in different modules, or there are just a lot of them, or it would be bad style to reference them from the parent class. Normal Java serialization does not have this problem (it just carries the actual dynamic class name with it), so this could be a barrier for adoption of Jackson for previous Java serialization users.

The problem here is that we want people to migrate from Java serialization. There are lots of reasons, most of them compelling.

Enter default typing. With default typing, we don’t need the type info annotations at all:

new ObjectMapper().enableDefaultTyping().writerFor(I.class).writeValueAsString(new A())
["net.hawo.tv.tvui17.A",{"i":0}]

(Ignore the fact that this is now suddenly an array – this is one of the ways jackson may include type info)
The idea is simple: Include the full class name in the json, and you can simply get the proper class at runtime! Sounds good, right?

The Vulnerability

Well, turns out this is not that good of an idea. Java serialization does a very similar thing (though it still requires the named class to be serializable, which jackson doesn’t) and this has lead to what feels like a third of all serious security vulnerabilities in Java applications, ever. The problem lies with the fact that Java classpaths are often massive, and allowing any class on that classpath to be deserialized at will can be disastrous since it exposes a huge attack surface. If you can get any class on the classpath to execute code when deserialized with jackson, you have successfully achieved remote code execution.
This is exactly what happened with Jackson. Some classes that were common on user classpaths could be deserialized to execute arbitrary code. The fix for this issue is basically a blacklist of a few of these classes that could be exploited. Blacklists are not a solution, though, and since this first list, the list has been amended several times. The maintainers are playing whack-a-mole here, and in my opinion it is a waste of everyone’s time to be adding all exploitable classes to this list.

The Solution

Our experience with this same issue in the Java serialization world tells us not to deserialize untrusted data. Luckily, Jackson is much more secure than Java serialization – if you don’t use default typing. The only acceptable solution to this issue in the long run is: do not use default typing to deserialize untrusted data. Default typing is rarely necessary or even a good idea.
Unfortunately, online resources saying this are sparse. Default typing is an “easy” solution, and many people simply do not have the security awareness to see the issue with it – they will stumble over stackoverflow answers such as this one and simply enable default typing to easily serialize Object fields. The documentation of Jackson also doesn’t highlight this as much as it should.
Two alternatives to default typing exist in Jackson:

Normal, annotation-based typing as shown above. This allows you either to use the full name as with default typing, or even specify your own name for greater compatibility (you can later change the class name without affecting the serialized representation). This is the “standard” solution, and the appropriate one for most use cases.
Should you not know the possible subtypes of a class you wish to deserialize in advance, you can use the rich registerSubtypes API to dynamically add the types you desire. These types could be detected through an existing module system you are already using, or using something like SPI.

All in all, I am a bit dissatisfied with the attention this issue has gotten. The issue is a security vulnerability by design, and anyone using default typing should have been aware of it. Luckily, default typing is not on by default. I do not have statistics but I would be surprised if many people used it or knew of its existence in the first place, and so I find the attention GitHub has given this a bit over the top – the biggest thing these notifications will spread is uncertainty about Jackson, so this article was an attempt at clearing up what it’s actually all about.

Gradle Properties

Gradle supports properties in builds just like Maven does, but the actual documentation and examples are harder to come by. This article will show you something that actually works.
In Maven, you typically have a dependencyManagement section that declares the dependencies with versions; I typically put the versions in a properties block so that they’re all located in a nice, handy, easy-to-manage place. See https://github.com/jottinger/ml/blob/master/pom.xml for a simple example; it’s not consistent with the property usage, but the idea should be clear. I have my versions all coalesced into an easy place to update them, and the changes propagate through the entire project.
Gradle examples rarely do anything like this. Therefore, for your edification, here’s a simple example:
gradle.properties:

aws_sdk_version=1.11.377
kotlin_version=1.2.50

build.gradle:

buildscript {
  repositories {
    mavenCentral()
  }
  dependencies {
    classpath "org.jetbrains.kotlin:kotlin-gradle-plugin:$kotlin_version"
  }
}
apply plugin: 'java'
apply plugin: 'kotlin'
repositories {
    mavenCentral()
    jcenter()
}
dependencies {
    compile "org.jetbrains.kotlin:kotlin-stdlib-jdk8"
    compile "com.amazonaws:aws-java-sdk:$aws_sdk_version"
}
tasks.withType(org.jetbrains.kotlin.gradle.tasks.KotlinCompile).all {
    kotlinOptions {
        jvmTarget = "1.8"
        javaParameters = true
    }
}

You can use a constraints block in the dependencies to do the same thing as dependencyManagement in Maven; that would look like this:

dependencies {
    constraints {
        implementation "com.amazonaws:aws-java-sdk:$aws_sdk_version"
    }
}

This would allow submodules (or this module) to use a simpler dependency declaration, leaving off the version, just like Maven can do:

dependencies {
    compile "org.jetbrains.kotlin:kotlin-stdlib-jdk8"
    compile "com.amazonaws:aws-java-sdk"
}

If you don’t want to use gradle.properties, that’s doable too (as pointed out by channel member matsurago) – you can put it in the buildscript block, as so:

buildscript {
  ext.aws_sdk_version="1.11.377"
}

Using multiple cores, the basics

The problem with multicore code

Having an application use more than one CPU during its execution introduces an explosion of execution paths. Instead of a simple “first Java executes the first line of this method, then, it executes the second line of this method, all the way to the end,” it becomes: “An arbitrarily chosen thread chooses an arbitrary number of instructions to execute, then the VM will make an arbitrary choice as to whether or not to make any changes to fields this thread performed visible to other threads unless you’ve done a good job on reading the Java memory model to ensure propagation of these changes.”
The former is easily understood. The latter is (largely) untestable and unfollowable, and therefore, it’s easy to write code that looks good, passes all tests, runs great on your machine… and bombs in production five days later, in a way that is utterly mystifying; no exception pointing at the problem. The problem won’t be reproducible (do the same thing, click the same buttons, feed in the same data… and now it works fine. And then tomorrow when that important customer logs in, it’ll fail again on you). Finding a bug like this is literally on the order of 100 to 1000 times harder to find, so avoiding even one such bug is ‘worth’ having 100 others.
Nevertheless, using all the cores in your CPU is:

Practical, in that just running it all on one can be unacceptably wasteful for performance, but more importantly
Inevitable, because you don’t write every line of code in your app yourself (you use libraries; the core java.* libraries at the very least), and THOSE will sometimes use threads.

So, what is a programmer to do?
You may feel like the only right answer is for the programmer to completely understand threading and learn how to write bug free code (given that tests can’t really find these bugs anyway, and even if you so happen to run into one, they are hard to reproduce and its hard to figure out which line(s) of code are faulty when you do find them). This means reading up on the entire Thread API and reading the Java Memory Model front-to-back (If thread 1 writes some data to some field, and sometime later thread 2 reads this field, it may or may not see the change. The JMM defines when Java is, and more importantly is not, required to have one thread ‘see’ updates of another.
But that’s a fool’s errand. You cannot guarantee writing bug free code.
Instead, anytime you have a need to have your code run on multiple cores, try to push towards going big or going small; in both cases, you NEVER call ANY method on Thread (possibly you do some Thread.sleep in the ‘going big’ style), you have no need whatsoever of synchronized, no state is shared between threads, and you don’t create new threads (some framework / library does this for you).

Going Big

Write or use a framework that runs your code for you and which takes care of multithreading. For example, a web framework is such a thing: You start it, and then it runs your code (and it takes care of creating whatever threads are necessary). Generally in these cases you do not write public static void main, or you do, but all that your main does is fire up the framework while you configure it or pass to it a bunch of ‘handlers’, and then the framework does the hard work for you.
Such frameworks are really good at setting up threading. This way each handler can be in its own little world; if a thread shares no data with any other, reasoning about threads is much, much simpler. Take webservers: Any given web handler simply does not get to interact with other handlers. It’s not like you can ask the web framework for a list of WebHandler objects that you can then call methods on.
There is often still a need to have 2 handlers interact. However, you should do this interaction by using tightly controlled communication channels where the issue of how concurrent operations interact is either irrelevant, or very well specified. There are 2 usual ways to go:

Databases. Databases define how concurrently running operations go via transactions. So, use transactions, make your database calls, and the database takes care of it. Communications-via-database is very common in web frameworks. You can use JDBI if you want to write SQL directly, or use Hibernate if you just want to store objects persistently.
Message queues, such as RabbitMQ. The idea behind these is that one handler will tell the message queue framework that it is interesting in all ‘foo’ events, and the other will tell the message queue framework: Here’s a ‘foo’ event, please deliver it to all handlers that said they’d like to know about it. All communication between handlers goes via the message queue.

Editor’s note: the projects mentioned here are very far from being your only choices. JDBI and Hibernate are good, but note that there are projects like myBatis, jOOQ, and others for SQL-ish libraries, and Hibernate is only one JPA implementation among many. Likewise for messaging: RabbitMQ is an AMQP 0.9 server, but you don’t use the JMS standard to talk to RabbitMQ; you do use JMS when using ActiveMQ, HornetQ, or any other of … quite a few messaging platforms. This is not saying that RabbitMQ or Hibernate, et al, are bad – your Editor uses them daily – but remember that your choices are not limited.

Going Small

Reduce the code that needs to run multithreaded to the smallest possible thing it can be, and then write a single stream/collection based operation that uses threading to apply this operation to a great many inputs concurrently. The fork/join framework is the usual go-to here.
Imagine you are writing a bitcoin miner. This operation can be described as follows:
GIVEN: A list of a few million randomly generated codes to try, and exactly 1 input block.
TASK: Write the hash into that 1 input block, hash it, and see if the hash ends up having the appropriate amount of 0s at the very end.
You can do this by going small: Write a trivial method (it won’t be larger than half a screen’s worth, common in this ‘go small’ model) which injects the code into the block, hashes it, and returns an empty string if the hash is not suitable, and if you hit the jackpot, returns the block. Then tell the framework you only want the non-empty returned values and voila, you’ve written a parallel bitcoin miner without ever having to touch java.lang.Thread.

Libraries

In practice, especially for the ‘going big’ route, you may need to have a cache or some such that should be shared between whatever threads your web framework is making for you, but try to find libraries for this, too.¹ There are some tricks to interacting with such libraries. Generally, you have to go ‘atomic’. For example, you can use the various collections in the java.util.concurrent package, but, the only guarantees it can make is that a single method call does the right thing. It cannot guarantee that a series of calls does the right thing. So, don’t do this:

if (concurrentMap.containsKey(myKey)) {
    String v = concurrentMap.get(myKey);
    operateOn(v);
}

Instead, do this:

String v = concurrentMap.get(myKey);
if (v != null) {
    operateOn(v);
}

The bad example can cause a problem: What if, in between the call concurrentMap.containsKey and the call to concurrentMap.get, some other thread removes the entry from the map? You’ll now call operateOn on a null value, not what was intended. Problematically, if this can happen, it will happen, but only very very rarely. Tests are unlikely to catch this problem.
Don’t do this:

String v = myCache.get(myKey);
if (v == null) {
    v = doExpensiveCalculationOfValue(myKey);
    myCache.put(myKey, v);
}

Instead, do this:

myCache.computeIfAbsent(myKey, k -> doExpensiveCalculationOfValue(k));

The reason to use computeIfAbsent here is because the first snippet can lead to running doExpensiveCalculationOfValue more than once. In the first snippet, imagine two (or more) threads get to the code with the same key about the same time. They’ll both find that there is no associated value (yet), so they both call doExpensibeCalculationOfValue and then they both set it. With computeIfAbsent, provided you use a properly concurrent map, such as java.util.concurrent.ConcurrentHashMap for example, doExpensiveCalculationOfValue is only ever called once, guaranteed.

Lessons

Use frameworks, either large: Web frameworks, or small: fork/join.
Thread-to-thread communication uses abstractions like message queues or databases. Don’t share state, don’t touch any fields from more than one thread.
If you must have direct thread-to-thread interop, use libraries with collections types that are designed for this, such as java.util.concurrent and guava‘s CacheBuilder stuff. When interacting with these collections, know that consecutive method calls to them have no guarantees of internal consistency, so reduce it to one call. Be aware of methods like java.util.Map‘s computeIfAbsent to enable this.

Footnotes

1 Editor’s note: caches are nontrivial in and of themselves. Go for “transactional caches” if you can – or just trade the speed that a cache would give you for the reliability that using a system of record gives you. Figure out what you need and do what fulfills that, before thinking that a cache is magic performance sauce – and note that the whole point of this article is that “magic performance sauce” usually isn’t what it says it is.

Article on java.time

Channel denizen (yes, that’s what we call people in the ##java channel, “denizens”) yawkat has published “An introduction to java.time,” an article to try to clear up some of the confusion around the time API in Java: things like LocalTime, OffsetTime, and the like. If you’re still using Date, this article is for you. Includes handy ways to convert between the time types, and is likely to grow further at need.

Caching All OkHttp Responses

When (integration) testing scrapers you need to strike a balance between “get as much input data as possible to cover parsing edge-cases” and “don’t DoS the backend, please”. Caching can help with that, and gives you a nice performance boost during testing on top.
The OkHttp http client library contains a cache built-in but by default it follows HTTP server caching headers. It also allows you to set a particular request to always use the cache, and error if the request isn’t present yet, but that is too harsh (we do want to fetch the first time after all). Turns out you can do both:

val cachedHttpClient = OkHttpClient.Builder()
        .cache(Cache(File(".download-cache/okhttp"), 512 * 1024 * 1024))
        .addInterceptor { chain ->
            chain.proceed(
                    chain.request().newBuilder()
                            .cacheControl(CacheControl.Builder()
                                    .maxStale(Integer.MAX_VALUE, TimeUnit.SECONDS)
                                    .build())
                            .build()
            )
        }
        .addNetworkInterceptor { chain ->
            log.info("Fetching {}", chain.request().toString())
            chain.proceed(chain.request())
        }
        .build()!!

This is kotlin code, but you get the point.

Modular password hashing with pwhash

When building web applications, we usually also have to store user authentication data. When doing so, there’s generally two choices: Either we use an external authentication provider like OAuth, or we store passwords for our users in our database. This blog post focuses on how to do the latter correctly.

Password hashing basics

Before we dive into the specifics of the pwhash library, we briefly discuss the basics of what we need to be wary of when storing passwords. The first, and most important point is that we do not actually ever store our users’ passwords. Instead, we run them through a cryptographic hash function. This way, in case our database ever gets compromised, the attacker doesn’t immediately know all our users’ passwords. But this is not considered enough any more. Because if we just put our users’ passwords through a simple hash function, like SHA-3, two users that pick the same password will end up with the same hash value. Knowing this is of course useful to an attacker, because now they can attack multiple passwords at the same time, if they get the user data.

Editor’s note: so choose passwords likely to be unique… and unguessable even to people who know you. Please.

Enter salts

To mitigate this problem, we use so-called salts. For each password, we generate a random salt, and prepend (or append) it to the password, before passing it through the hash function. This leaves us with
sha3(firstsalt:password) = 0bee3940e2d74f5155e73a9e90ea75b5d06407db85527f62acefa97af2d59f69589b276630e20a1ff8b0b781a372ae17db88b9f782acf7ed0022ab4c2fc766df
sha3(secondsalt:password) = edaa1e8b31e2d2943d689928a5ba1c503bb3220ddc164d6feff9e178dd18a4727b1905400252ef071d28b370c0b9727420cd0e2011109ca3e8934a4082d2bf9e
As we can see, this yields two completely different hashes even though two users used the same (admittedly very bad) password. The salt can be stored alongside the hash in the database. An attacker now has to break each password individually, instead of attacking them all at the same time. Great, right? Surely, now we’re done and can get to the actual library? Not quite yet. The next problem we have is that modern graphics cards are really fast at calculating these kind of hashes. For example, an Nvidia GTX 1080 calculates around 800 million SHA-3 hashes per second.

Dedicated password hashing functions

To get rid of this problem, we do something we usually don’t want to do in computing: We make things intentionally slow. There are several functions that can be used for this. An older idea is simply applying many rounds of the same hash function over and over again. An example of this approach is the PBKDF2 (Password-Based-Key-Derivation-Function). But since GPUs are really fast these days, we also want functions that are harder to compute on GPUs. Functions that use a lot of memory, and a lot of branches are generally very hard to compute on graphics cards. An example of this is argon2, a function which was specifically designed for hashing passwords, and won the Password Hashing Competition in 2015. The question, then, is how do we create a re-usable approach that allows us to stay current with always-current password hashing requirements?

The pwhash library

While there are already Java implementations and bindings for argon2 and other password hashing functions such as bcrypt, what is still missing in the Java ecosystem is a library that unifies them under a single interface. Functions like argon2 come with different versions, and a lot of adjustable parameters. So if we hardcode those parameters into our application in 2018, the parameters are probably going to be outdated (and inefficient or insecure) in five or ten years. So ideally, we want our library to handle this for us. We change our parameters in one place, and the library automatically takes care of upgrading both new and existing password hashes every time a user logs in or signs up. In addition to this, it would also be convenient to not just switch between parameters of a single algorithm, but also to switch the algorithm to something newer and better altogether.
That’s the goal of the pwhash library.

The HashStrategy interface

To offer all this, the core functionality of pwhash is the HashStrategy interface. It offers 3 methods:

String hash(String password)
boolean verify(String password, String hash)
boolean needsRehash(String password, String hash)

This interface is largely inspired by the PHP APIs for password hashing, which offers the functions password_hash, password_verify and password_needs_rehash. In the author’s opinion, this is one of the better things in the PHP core library, and hence I decided to port the functionality to Java in a bit more Object-Oriented style.
The first two methods are relatively straightforward: String hash(myPassword) produces a hash, alongside with all its parameters, which can later be passed to boolean verify(myPassword, storedHash) in order to verify if the password actually matches the given hash.
The third method is the magic that lets us upgrade our hashes without needing to write new code every time. In the first step, it checks whether the given password actually verifies against the hash, in order to make sure that we don’t accidentally overwrite hashes when the user supplied an incorrect password. In the next step, the parameters used for the current HashStrategy (which are generally supplied by the constructor, or some factory method) are compared with the ones stored alongside the password hash. If they match, false is returned, and no further action needs to be taken. If they do not match, the method returns true instead. Now, we call hash(myPassword) again, and store the newly generated hash in the database.
This way, we only have to write code once for our login, and every time we change our parameters, they are automatically updated every time users log in. However, for a single implementation, like the Argon2Strategy, this only handles changes in parameters. If somehow a weakness is discovered, we still do not have a good way to migrate away from Argon2 to a different algorithm.

The MigrationStrategy class

Migration between two different algorithms is handled by the MigrationStrategy class, which implements the above interface, and in its constructor, accepts two other strategies: One to migrate from, and another to migrate to. Its implementation for String hash(String) simply calls the hash functions of the latter strategy, since we want all new hashes to be performed with the new algorithm.
Its implementation of boolean verify(String, String) first attempts to verify against the old strategy, and if that fails, against the new strategy. If neither succeed, the password was incorrect. If either succeeds, the password was correct, and true can be returned.
Again, the boolean needsRehash(String, String) method is a little bit more complicated. It first attempts to verify the given password against the old strategy. If this succeeds, the password is obviously in the old format, and needs to be rehashed, and we immediately return true. In the second step, we attempt to verify it with the new strategy. If this is also unsuccessful, we return false as we don’t want to rehash if the user supplied an incorrect password. If the verification succeeds, we return whatever newStrategy.needsRehash(String, String) returns.
This way, we can adjust parameters on our new strategy while some passwords are still hashed with the old strategy, and we receive the expected results.
Note that it is not necessary to use this class when migrating parameters inside one implementation. It is only used when switching between different classes.

Chaining MigrationStrategy

Fun fact about MigrationStrategy: Since it is also an implementation of the HashStrategy interface, you can even nest it in another MigrationStrategy like so:

MigrationStrategy one = new MigrationStrategy(veryOldStrategy, somewhatBetterStrategy);
MigrationStrategy two = new MigrationStrategy(one, reallyGoodStrategy);

This might not seem useful at first, but it is useful if you have a lot of users. You may want to switch to yet another strategy at some point, while not all your users are migrated to the intermediate strategy yet. Chaining the migrations like this allows you to successfully verify against all three strategies, still allowing logins for even your oldest users who haven’t logged in in a while, and still migrating all passwords to the newest possible algorithm.

Examples

Examples for the usage of the pwhash library are hosted in the repository on GitHub.
Currently, there are two examples. One uses only the Argon2Strategy and upgrades parameters within that implementation. The second example uses the MigrationStrategy to show code that migrates users from a legacy Pbkdf2Strategy to a more modern Argon2Strategy.
The best thing about these examples: The code that actually authenticates users is exactly the same in both examples. The only thing that is different is the HashStrategy that is plugged in. This is the main selling point of the pwhash library: You write your authentication code once, and when you want to upgrade, all you have to do is plug in a new strategy.

Where to get it

pwhash is available on Maven Central. For new applications, it is recommended to only use the -core artifact. Support for PBKDF2 is mainly targeted at legacy code bases wanting to migrate away from PBKDF2. JavaDoc is available online on the project homepage. JAR archives are also available on GitHub, but it is strongly recommended you use maven instead.