we are using clj-otel 0.2.10 with honeycomb in-cloud and jaeger-all-in-one for local dev. in cloud, it works great; spans come through to HC just fine. in local, nothing shows up in jaeger. it used to work, but somewhere in the past few months, it stopped working. if i switch local to honeycomb, that works. jaeger-all-in-one is functional in that its ui is displaying its own ui's traces, however, nothing comes through from our app. we're using all the default settings - ports etc. my question -- how do i go about debugging this from the clj-otel side? GPT gave me a bunch of java opts and logback settings to enable but they've produced nothing new to observe (and so may just be irrelevant slop).
Update for the curious: we were emitting spans before our otel sdk was started, putting the system into an invalid state due to how the tracer is cached, as illustrated by this exception
Caused by: java.lang.IllegalStateException: GlobalOpenTelemetry.set has already been called. GlobalOpenTelemetry.set must be called only once before any calls to GlobalOpenTelemetry.get. If you are using the OpenTelemetrySdk, use OpenTelemetrySdkBuilder.buildAndRegisterGlobal instead. Previous invocation set to cause of this exception.
at io.opentelemetry.api.GlobalOpenTelemetry.set(GlobalOpenTelemetry.java:107)Many of the settings you list above do not exist. You can review the documentation for autoconfigure properties here: https://opentelemetry.io/docs/languages/java/configuration/#environment-variables-and-system-properties
For your local setup, I suggest configuring the application to send telemetry to a Collector instance. In the Collector config, you could add the debug exporter, and ensure each pipeline includes the debug exporter. The Collector will then output a brief summary of signals received. Here is an example : https://github.com/steffan-westcott/clj-otel/blob/master/examples/rpg-service/otel-collector.yaml
thank you, as i expected 😄 i'll try that
Are you using the instrumentation agent? If so, be aware a breaking change was introduced with the 2.x versions of the agent, in that the OTLP defaults changed from grpc to http/protobuf, with associated changes in port numbers. All the clj-otel examples that include the agent assume 2.x is used. Note only the defaults for the agent are affected - The defaults when using the SDK remain as grpc.
i will look into this
i guess a quick test is to ensure that the http/protobuff ports are enabled on jaeger. if it works, that'd be the cause
possibly ai slop, but sharing nonetheless logback settings
<logger name="com.steffanwestcott" level="DEBUG"/>
<logger name="io.opentelemetry" level="DEBUG"/>
a deps.edn alias with the following jvm opts:
[ ;; ===== OpenTelemetry core logging =====
"-Dio.opentelemetry.logger.level=DEBUG"
;; ===== OTLP exporter deep debugging =====
"-Dotel.java.experimental.exporter.debug=true"
"-Dotel.java.experimental.exporter.otlp.retry.verbose=true"
;; ===== Force OTLP + endpoints =====
"-Dotel.traces.exporter=otlp"
"-Dotel.exporter.otlp.endpoint="
"-Dotel.exporter.otlp.protocol=grpc"
;; ===== Make flushing & metrics faster =====
"-Dotel.bsp.schedule.delay=100"
"-Dotel.metric.export.interval=1000"
;; ===== Resource visibility =====
"-Dotel.resource.attributes=service.name=my-service,env=local"
;; ===== Disable sampling so NOTHING is dropped =====
"-Dotel.traces.sampler=always_on"] we are not using the instrumentation agent, just the SDK